---
title: "Slides: The Art & Science of Benchmarking Agents — Vincent Chen, Snorkel AI"
category: "slides"
video_id: "iNkFlCiij0U"
sourceLabels: ["Public YouTube video frames", "Public YouTube metadata"]
---

# Slides: The Art & Science of Benchmarking Agents — Vincent Chen, Snorkel AI

## Source Video
[The Art & Science of Benchmarking Agents — Vincent Chen, Snorkel AI](https://www.youtube.com/watch?v=iNkFlCiij0U)

## Relationship To World's Fair 2026
These slides are extracted from a public AI Engineer YouTube video connected to World's Fair 2026. Speaker-matched clips are supporting context unless later confirmed as exact session recordings; official livestream recordings are day-level/event-level source material.

## Related Scheduled Sessions
- No individual scheduled session mapping has been assigned yet; treat this as an event livestream deck.

## Extracted Slides
![[assets/slides/iNkFlCiij0U/slide-002.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/iNkFlCiij0U/slide-002.html)
- AI slide classifier: `content_slide` confidence `0.97`
- Text source: advanced OCR `rapidocr-live/right-72/opencv-adaptive`.
- OCR decision: ready — dense multi-column slide with small text and logos

Slide text:

> Snorke The Frontier Al Data Lab
> Academic Labs! Team Publications/OsS Collsborations
> Alex Ratrer Stmlord Ppo. Co-TO LCEO Asu, Prol Ltult iLlL anduyry Jnd lesdung instiution moruing on dutu (enlrk A mith fexng global Lbs Rtsearchers, FDEs, and Applied A Erginrers frorn. ondLLTenkriAnd foundlnon modthy
> Chris Re UA
> Prof. at Susfeeu Co foundir. Vanfotd Center for Hretrch on 1 Yalc A+urd at Hrurts, Ktl, klk, Ull, Y.Dt, an4 more
> Expertlso and Capablrties Bonchmarks
> Frod Sats UCLA PhD.Aust Prof.st ClirSnolan Hieh prror proramTk ala dreapment lor Cotachartd + cetihutd to Uo Tnrb.ein. Liart-Trr tt et adancd phyk1 proolrs
> Aerrth aorkhoa, harrs, ard toct u+ rmcdetn. -
> 12
> Bnere A I Popr tlly + Carld+ren t Not yr dHttuler
> Google DeepMind

![[assets/slides/iNkFlCiij0U/slide-003.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/iNkFlCiij0U/slide-003.html)
- AI slide classifier: `content_slide` confidence `0.97`
- Text source: advanced OCR `rapidocr-live/full`.
- OCR decision: ready — simple title slide with prominent text plus sponsor logos

Slide text:

> Open Benchmarks
> AIE Grants
> A$3M+commitmentto
> openbenchmarks
> Snorkel HuggingFace PRImEIntellect together.ai FACTORY harbor OPyTorch
> AIEngineer
> EUROPE

![[assets/slides/iNkFlCiij0U/slide-004.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/iNkFlCiij0U/slide-004.html)
- AI slide classifier: `content_slide` confidence `0.98`
- Text source: advanced OCR `rapidocr-live/right-72/contrast`.
- OCR decision: ready — dense slide with chart panels and small text

Slide text:

> Difficulty / Model Headroom
> Benchmarkis unsaturated.
> ARC-AGI-3
> tasks were solvable by humans) Shipped with <1% model score (100% of
> The Human-Al Gap
> Engineering the future of Al

![[assets/slides/iNkFlCiij0U/slide-005.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/iNkFlCiij0U/slide-005.html)
- AI slide classifier: `title_card` confidence `0.94`
- Text source: agent_vision.

Slide text:

> Lasting benchmarks push the frontier


### Hidden Non-Slide Evidence
- [`slide-001.jpg`](/assets/slides/iNkFlCiij0U/slide-001.jpg) — `speaker_stage` confidence `0.96`; camera shot of speaker and audience with projected slide; not a clean presentation slide frame
- [`slide-006.jpg`](/assets/slides/iNkFlCiij0U/slide-006.jpg) — `speaker_stage` confidence `0.95`; camera shot of speaker and audience with projected slide; not a clean presentation slide frame
- [`slide-007.jpg`](/assets/slides/iNkFlCiij0U/slide-007.jpg) — `title_card` confidence `0.98`; Branding/title card with logo and URL only; no substantive presentation content.

Classification audit: `raw/sources/slide-ai-classification/slides/iNkFlCiij0U/audit.json`

## Slide-Derived Subjects To Review
Subject extraction uses video title, related session titles/descriptions, transcript context, and OCR text when available. OCR is best-effort and should be reviewed against the embedded slide images.
