Slides: Voice In, Visuals Out: The Agony and the Ecstasy - Allen Pike, Forestwalk Labs
Source Video
Voice In, Visuals Out: The Agony and the Ecstasy - Allen Pike, Forestwalk Labs
Relationship To World's Fair 2026
These slides are extracted from a public AI Engineer YouTube video connected to World's Fair 2026. Speaker-matched clips are supporting context unless later confirmed as exact session recordings; official livestream recordings are day-level/event-level source material.
Related Scheduled Sessions
- No individual scheduled session mapping has been assigned yet; treat this as an event livestream deck.
Extracted Slides

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
title_cardconfidence0.99 - Text source: agent_vision.
Slide text:
Voice In, Visuals Out
The Agony and the Ecstasy

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.98 - Text source: agent_vision.
Slide text:
“Audio is the human-preferred input to AIs, but vision is the preferred output from them.”
— Andrej Karpathy, May 2026

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.97 - Text source: agent_vision.
- OCR decision: ready — Dense diagram and UI screenshot text are better handled by OCR; only short obvious labels are captured here.
Slide text:
Visualization
Interactivity

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.97 - Text source: agent_vision.
- OCR decision: ready — Dense diagram/UI content with small labels is OCR-suitable; only the short visible labels are captured here.
Slide text:
Visualization
Interactivity
Beauty

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
title_cardconfidence0.95 - Text source: agent_vision.
Slide text:
Voice In

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
demo_videoconfidence0.94 - Text source: none.
- OCR decision: ready — The slide is a row of video thumbnails with small captions and view counts; OCR is better suited than manual transcription in this pass.
- Slide text: not surfaced (
illegibleby AI classifier).

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
title_cardconfidence0.98 - Text source: agent_vision.
Slide text:
The Tyranny of Latency

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.98 - Text source: agent_vision.
Slide text:
200ms seamless voice
100ms feels instant

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.99 - Text source: agent_vision.
Slide text:
Pillars of low latency
1 Fast models

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.99 - Text source: agent_vision.
Slide text:
Pillars of low latency
1 Fast models
2 Short intervals
3 Stable cache
Hidden Non-Slide Evidence
- `slide-007.jpg` —
sponsor_logoconfidence0.98; logo-only branding frame with no substantive slide content
Classification audit: raw/sources/slide-ai-classification/slides/65X0pQ6Lmbg/audit.json
Slide-Derived Subjects To Review
Subject extraction uses video title, related session titles/descriptions, transcript context, and OCR text when available. OCR is best-effort and should be reviewed against the embedded slide images.