---
title: "From Zero to Leaderboard: Building an End-to-End AI Agent Evaluation Pipeline"
category: "talks"
date: "2026-06-29"
time: "12:10pm-1:10pm"
track: "Workshops Day 1"
room: "Track 5"
speakers: ["Wolfram Ravenwolf"]
sourceLabels: ["Official conference schedule", "Public YouTube metadata"]
scheduleTrack: "Workshops Day 1"
scheduleRoom: "Track 5"
scheduleLabels: ["Workshops Day 1", "Track 5", "workshop", "confirmed"]
---
# From Zero to Leaderboard: Building an End-to-End AI Agent Evaluation Pipeline

## Conference Context
- Date/time: 2026-06-29 · 12:10pm-1:10pm
- Track/room: Workshops Day 1 · Track 5
- Speaker(s): Wolfram Ravenwolf
- Session type/status: workshop · confirmed

- Track: Workshops Day 1
- Room: Track 5
- Session type: workshop
- Status: confirmed

## Session Description
Running one agent eval is easy. Running hundreds — with controlled timeouts, replicated configs, and automated collection across distributed VMs — requires infrastructure that most teams end up building from scratch. In this workshop, we shortcut that process and build a rigorous evaluation pipeline end-to-end. Participants will set up and connect the full evaluation stack: **Layer 1 — The Benchmark Runner.** Configure Harbor to orchestrate parallel agent evaluations on Terminal-Bench 2.0, with W&B Sandboxes providing isolated environments for each task. **Layer 2 — The Collection Pipeline.** Use WolfBench to scan distributed VMs for results, deduplicate across runs, download trajectories, and build a local results archive that survives VM teardown. **Layer 3 — The Analysis Framework.** Compute the five-metric framework (Ceiling / Best / Average / Worst / Solid) across replicated runs. Learn to read the spread: when is a model "better"? When is a score difference just noise? **Layer 4 — The Observability Layer.** Upload full agent conversation traces to W&B Weave for per-turn inspection. See exactly where an agent goes wrong — the command it ran, the output it misread, the moment it started looping. **Layer 5 — The Leaderboard.** Generate interactive HTML charts that show the full performance distribution, not a single bar. We'll work with real data from hundreds of production runs, and participants will leave with a working pipeline they can adapt to their own agents and benchmarks. Laptops required; all tools are open-source.

## Media Evidence
No related AI Engineer channel video found yet.

## Evidence Graph
This evidence graph is generated from currently linked source material: official schedule text, related video pages, cached transcripts, visible slide text, dense/reconstructed slide pages, and AI slide-classification audits.

### Media Signals
No linked video, transcript, or slide source has been attached yet.

### Agent Reading Notes
Use these signals to refine the synopsis, topic links, people/company context, and method notes. If a source is a related external video rather than an exact official recording, keep it framed as supporting evidence.

## Transcript Status
No official session recording transcript was found by exact title match on the AI Engineer YouTube channel during this run.

## People
- [[wolfram-ravenwolf]]

## Notes
- Pending transcript synthesis when an official recording or confirmed matching video is available.

## Synthesis
### Synthesized Breakdown
# From Zero to Leaderboard: Building an End-to-End AI Agent Evaluation Pipeline ## Conference Context - Date/time: 2026-06-29 · 12:10pm-1:10pm - Track/room: Workshops Day 1 · Track 5 - Speaker(s): Wolfram Ravenwolf - Session type/status: workshop · confirmed - Track: Workshops Day 1 - Room: Track 5 - Session type: workshop - Status: confirmed ## Session Description Running one agent eval is easy. Running hundreds — with controlled timeouts, replicated configs, and automated collection across distributed VMs — requires infrastructure that most teams end up building from scratch. In this workshop, we shortcut that process and build a rigorous evaluation pipeline end-to-end. Participants will set up and connect the full evaluation stack: **Layer 1 — The Benchmark Runner.** Configure Harbor to orchestrate parallel agent evaluations on Terminal-Bench 2.0, with W&B Sandboxes providing isolated environments for each task.

### Speaker And Company Context
- [[wolfram-ravenwolf|Wolfram Ravenwolf]] — AI Evangelist at [[weights-and-biases-by-coreweave|Weights & Biases by CoreWeave]].

### Topics Covered
- [[agent-security]]
- [[ai-sandboxes]]

### Derived Links And Source Material

### Novel Concepts / Clever Methods
- No highlighted novel concept has been detected yet.

### Evidence Boundary
This synthesis is based on the official schedule and linked source pages. It should be revisited when exact session recordings or transcript-backed secondary sources are available.
