---
title: "Slides: Production Evals For Agentic AI Systems - Nishant Gupta, Meta Superintelligence Labs"
category: "slides"
video_id: "vljxQZfJ9wY"
sourceLabels: ["Public YouTube video frames", "Public YouTube metadata"]
---

# Slides: Production Evals For Agentic AI Systems - Nishant Gupta, Meta Superintelligence Labs

## Source Video
[Production Evals For Agentic AI Systems - Nishant Gupta, Meta Superintelligence Labs](https://www.youtube.com/watch?v=vljxQZfJ9wY)

## Relationship To World's Fair 2026
These slides are extracted from a public AI Engineer YouTube video connected to World's Fair 2026. Speaker-matched clips are supporting context unless later confirmed as exact session recordings; official livestream recordings are day-level/event-level source material.

## Related Scheduled Sessions
- No individual scheduled session mapping has been assigned yet; treat this as an event livestream deck.

## Extracted Slides
![[assets/slides/vljxQZfJ9wY/slide-001.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/vljxQZfJ9wY/slide-001.html)
- AI slide classifier: `title_card` confidence `0.99`
- Text source: agent_vision.

Slide text:

> Production Evals
> for Agentic Systems
> Measuring reliability beyond accuracy. Building evaluation systems for autonomous AI workflows.
> Nishant Gupta
> Tech Lead @ Meta

![[assets/slides/vljxQZfJ9wY/slide-002.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/vljxQZfJ9wY/slide-002.html)
- AI slide classifier: `content_slide` confidence `0.98`
- Text source: advanced OCR `rapidocr-live/bright-screen/contrast`.
- OCR decision: ready — Small chart labels, panel headers, and callouts are OCR-suitable.

Slide text:

> our evaluation methods AI systems evolved faster than
> The Illusion The Reality
> 100% Modes Invisible Failure
> 75% Behavior Degraded Production
> Benchmark Accuracy 90% 25% 50% 0% T-0 T+10ms T+50ms Reliability Gaps Unpredictable User MAAAA T+10GmS

![[assets/slides/vljxQZfJ9wY/slide-003.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/vljxQZfJ9wY/slide-003.html)
- AI slide classifier: `content_slide` confidence `0.98`
- Text source: advanced OCR `rapidocr-live/border-trim/contrast`.
- OCR decision: ready — Structured table text and cell labels are OCR-suitable.

Slide text:

> The Paradigm Shift: Output vs. Behavior
> Traditional LLM Evaluation Agent Evaluation
> Goal Output Accuracy Workflow Behavior
> Environment Static Datasets Dynamic Contexts
> Execution Single-path Processing Multi-path & Tool Dependent
> Failure Mode Hallucination Cascading Workflow Failure

![[assets/slides/vljxQZfJ9wY/slide-004.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/vljxQZfJ9wY/slide-004.html)
- AI slide classifier: `content_slide` confidence `0.94`
- Text source: agent_vision.

Slide text:

> Think like an SRE: Accuracy gives way to Reliab

![[assets/slides/vljxQZfJ9wY/slide-005.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/vljxQZfJ9wY/slide-005.html)
- AI slide classifier: `content_slide` confidence `0.98`
- Text source: agent_vision.

Slide text:

> The Evaluation Signal Hierarchy
> Production Telemetry
> Scenario Evals
> Benchmarks

![[assets/slides/vljxQZfJ9wY/slide-006.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/vljxQZfJ9wY/slide-006.html)
- AI slide classifier: `content_slide` confidence `0.98`
- Text source: advanced OCR `rapidocr-live/border-trim/opencv-adaptive`.
- OCR decision: ready — Small diagram labels and the metrics box are OCR-suitable.

Slide text:

> Offline Evals: Scenario-Driven Simulation
> Agent Sandbox Discrete Outputs
> Test Runner Sinulated. Tools Execution: Steps Update State Hetrics Conpletion Rate 98.5x
> .Tool Correctness 108%
> Plan Quality High'
> Simulated Cost $0.05
> Scenario-driven, not prompt-driven.

![[assets/slides/vljxQZfJ9wY/slide-007.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/vljxQZfJ9wY/slide-007.html)
- AI slide classifier: `content_slide` confidence `0.97`
- Text source: agent_vision.

Slide text:

> Online Evals: The Production Stream Production is your largest evaluation dataset. Every interaction is signal.

![[assets/slides/vljxQZfJ9wY/slide-008.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/vljxQZfJ9wY/slide-008.html)
- AI slide classifier: `content_slide` confidence `0.97`
- Text source: agent_vision.

Slide text:

> Human-in-the-Loop Calibration Humans are evaluators, not merely fallback systems.

![[assets/slides/vljxQZfJ9wY/slide-009.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/vljxQZfJ9wY/slide-009.html)
- AI slide classifier: `content_slide` confidence `0.98`
- Text source: advanced OCR `rapidocr-live/border-trim/opencv-adaptive`.
- OCR decision: ready — dense chart and dashboard text are better suited for OCR

Slide text:

> Observability is the Prerequisite
> The Trace Waterfall Live Metrics Dashboard.
> User Prompt 3sm-70ms) Latency 345 ms
> Planner lteration (7a-3sms)
> Retries 7
> Vector DB Lookup. 8ms 2.5 % AW
> Parallel APl Tool Calls.- AP1B:38ms APIA-45m) Step Costs $0.014
> APIC: 62mu Memory Usage
> 480 MB
> "You cannot evaluate what you cannot observe.

![[assets/slides/vljxQZfJ9wY/slide-010.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/vljxQZfJ9wY/slide-010.html)
- AI slide classifier: `content_slide` confidence `0.97`
- Text source: agent_vision.

Slide text:

> The Continuous Evaluation Loop Evaluation is an always-running service, not a testing phase.

![[assets/slides/vljxQZfJ9wY/slide-011.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/vljxQZfJ9wY/slide-011.html)
- AI slide classifier: `content_slide` confidence `0.96`
- Text source: agent_vision.

Slide text:

> The Agentic Control Plane Reference Architecture

![[assets/slides/vljxQZfJ9wY/slide-012.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/slides/vljxQZfJ9wY/slide-012.html)
- AI slide classifier: `content_slide` confidence `0.99`
- Text source: agent_vision.

Slide text:

> Architectural Imperatives
> 1. Offline benchmarks are necessary but insufficient.
> 2. Agentic systems must be evaluated as full workflows.
> 3. Production telemetry is the ultimate evaluation signal.
> 4. Reliability always supersedes raw model accuracy.
> 5. Evals are no longer tests; they are core infrastructure.
> You can't improve what you don't continuously evaluate.


Classification audit: `raw/sources/slide-ai-classification/slides/vljxQZfJ9wY/audit.json`

## Slide-Derived Subjects To Review
Subject extraction uses video title, related session titles/descriptions, transcript context, and OCR text when available. OCR is best-effort and should be reviewed against the embedded slide images.
