---
title: "How long can your skills be before your agent forgets what you told it?"
category: "talks"
date: "2026-06-30"
time: "1:30pm-1:50pm"
track: "Context Engineering"
room: "Track 8"
speakers: ["Laurie Voss"]
sourceLabels: ["Official conference schedule", "Public YouTube metadata"]
scheduleTrack: "Context Engineering"
scheduleRoom: "Track 8"
scheduleLabels: ["Context Engineering", "Track 8", "session", "confirmed"]
---
# How long can your skills be before your agent forgets what you told it?

## Conference Context
- Date/time: 2026-06-30 · 1:30pm-1:50pm
- Track/room: Context Engineering · Track 8
- Speaker(s): Laurie Voss
- Session type/status: session · confirmed

- Track: Context Engineering
- Room: Track 8
- Session type: session
- Status: confirmed

## Session Description
A year ago, frontier models lost the thread somewhere around 200 simultaneous instructions, so skills files had to stay short and lean on sub-skills and subagents. We re-ran IFScale on the 2026 frontier and found the ceiling has moved by an order of magnitude: closer to 2,000 instructions, up to 5,000 on the strongest models. The more interesting story is how models fail at the new frontier: DeepSeek quietly drops instructions, Opus refuses outright when innocuous words trip a safety classifier, Gemini burns its whole budget on reasoning and emits nothing, and GPT-5.5 stops to tell you your request was unreasonable. The capacity problem is largely solved; verification is wide open. We'll show the data, the failure modes, and what it costs to find out. You’ll come out with hard data on the ceiling for complex instructions to LLMs

## Media Evidence
[Ship Real Agents: Hands-On Evals for Agentic Applications — Laurie Voss, Arize](https://www.youtube.com/watch?v=Xfl50508LZM) (speaker-match related prior/adjacent AI Engineer video; captions: English auto-captions).

- [[youtube-Xfl50508LZM-transcript]] — full cached transcript markdown for the related YouTube source.

- Source video: `youtube-Xfl50508LZM`
- Slide deck: [[youtube-Xfl50508LZM-dense-slides|Dense Slides: Ship Real Agents: Hands-On Evals for Agentic Applications — Laurie Voss, Arize]] — 7 visible slide image(s); 7 HTML recreation(s).
![[assets/dense-slides/Xfl50508LZM/slide-001.jpg]]
![[assets/dense-slides/Xfl50508LZM/slide-002.jpg]]
![[assets/dense-slides/Xfl50508LZM/slide-003.jpg]]
- Additional slide evidence: [[youtube-Xfl50508LZM-slides|Slides: Ship Real Agents: Hands-On Evals for Agentic Applications — Laurie Voss, Arize]], [[youtube-Xfl50508LZM-reconstructed-slides|Reconstructed Slides: Ship Real Agents: Hands-On Evals for Agentic Applications — Laurie Voss, Arize]]
- Slide-derived themes for `youtube-Xfl50508LZM`: phoenix, prompt, settings, general, detect, regressions, change, compare.

## Evidence Graph
This evidence graph is generated from currently linked source material: official schedule text, related video pages, cached transcripts, visible slide text, dense/reconstructed slide pages, and AI slide-classification audits.

### Media Signals
- `youtube-Xfl50508LZM` — 22,591 transcript words; 6 slide-derived text signals
- Transcript signals for `youtube-Xfl50508LZM`: evals, eval, data, should, judge, output, whether, phoenix.
- Slide-derived themes for `youtube-Xfl50508LZM`: phoenix, prompt, settings, general, detect, regressions, change, compare.
- Evidence links for `youtube-Xfl50508LZM`: [[youtube-Xfl50508LZM]], [[youtube-Xfl50508LZM-transcript]], [[youtube-Xfl50508LZM-slides]], [[youtube-Xfl50508LZM-dense-slides]], [[youtube-Xfl50508LZM-reconstructed-slides]]

### Agent Reading Notes
Use these signals to refine the synopsis, topic links, people/company context, and method notes. If a source is a related external video rather than an exact official recording, keep it framed as supporting evidence.

## Transcript Status
Related video transcript availability: English auto-captions. Treat this as supporting context, not a recording of this exact scheduled session unless later confirmed. Cached at `raw/sources/youtube-transcripts/Xfl50508LZM.txt` (22,591 words).

## People
- [[laurie-voss]]

## Supporting Slides
- [[youtube-Xfl50508LZM-slides]] — extracted from the related public AI Engineer video.

## Slide Evidence
- Slide-only cropped deck: [[youtube-Xfl50508LZM-dense-slides]] (7 viable slide images).
- Related slide/OCR pages:
- [[youtube-Xfl50508LZM-dense-slides]]
- [[youtube-Xfl50508LZM-reconstructed-slides]]
- [[youtube-Xfl50508LZM-slides]]
- Slide-derived terms: `phoenix`, `claude`, `tome`, `setting`, `tracing`, `alengineer`, `europe`, `ages`, `notebook`, `cloud`, `comma`, `swiss`, `cheese`, `braintrust`, `workos`, `openal`, `frage`, `gers`

## Synthesis
### Synthesized Breakdown
Hi everybody. Uh my name's Laurie Voss. I am head of developer experience at Arize AI. Uh in a former life, I co-founded npm Inc.

### Speaker And Company Context
- [[laurie-voss|Laurie Voss]] — Head of Developer Relations at [[arize-ai|Arize AI]].

### Topics Covered
- [[agent-security]]
- [[agentic-search]]
- [[agentic-web]]
- [[ai-sandboxes]]
- [[coding-agents]]

### Derived Links And Source Material
- [[youtube-Xfl50508LZM-transcript]] — transcript markdown; source cache `raw/sources/youtube-transcripts/Xfl50508LZM.txt` (22,591 words).
- [[youtube-Xfl50508LZM]] — related YouTube source page.
- [[youtube-Xfl50508LZM-slides]] — slide evidence.
- [[youtube-Xfl50508LZM-reconstructed-slides]] — slide evidence.
- [[youtube-Xfl50508LZM-dense-slides]] — slide evidence.

### Novel Concepts / Clever Methods
- No highlighted novel concept has been detected yet.

### Evidence Boundary
This synthesis uses the official schedule plus cached video transcripts. Official AI Engineer World's Fair San Francisco 2026 livestreams and cut videos are primary event video sources for transcript/slide evidence; external, historical, or speaker-matched videos remain supporting context unless manually verified as exact official event recordings.
