2 hr deep dive on LLM Inference at Scale — Part 2 of 2
Conference Context
- Date/time: 2026-06-29 · 1:15pm-2:15pm
- Track/room: Workshops Day 1 · Track 3
- Speaker(s): Harshul Jain, Tanmay Sah
- Session type/status: sponsor · confirmed
- Track: Workshops Day 1
- Room: Track 3
- Session type: sponsor
- Status: confirmed
Session Description
Most engineers using LLMs can call an API. Far fewer can explain why their model is slow, why it's running out of memory, or how the inference engines powering every major LLM API actually work. This workshop walks through the full inference stack — from how a transformer generates a single token to serving billions of tokens a day with vLLM, SGLang, TensorRT-LLM, Ray, and KServe/llm-d. 60% explanation with live demos, 40% hands-on exercises. Attendees leave with a running vLLM server they benchmarked themselves. Based on the open-source practitioners handbook being built live at github.com/harshuljain13/llm-inference-at-scale (NOTE: this is a 2 hour workshop that happens over lunch break - you should try to have lunch before or after if attending)
Media Evidence
No related AI Engineer channel video found yet.
These are phone-photo slide captures from the Google Photos AIE Slides album. They are supporting slide evidence and do not override official schedule fields.
- google photos aie slides 9gWZzS1EpXM1C5eK6 llm inference at scale slides - Google Photos Slides: LLM Inference at Scale Workshop (confidence: high).
Evidence Graph
This evidence graph is generated from currently linked source material: official schedule text, related video pages, cached transcripts, visible slide text, dense/reconstructed slide pages, and AI slide-classification audits.
Media Signals
No linked video, transcript, or slide source has been attached yet.
Agent Reading Notes
Use these signals to refine the synopsis, topic links, people/company context, and method notes. If a source is a related external video rather than an exact official recording, keep it framed as supporting evidence.
Transcript Status
No official session recording transcript was found by exact title match on the AI Engineer YouTube channel during this run.
People
Notes
- Pending transcript synthesis when an official recording or confirmed matching video is available.
Synthesis
Synthesized Breakdown
2 hr deep dive on LLM Inference at Scale — Part 2 of 2 ## Conference Context - Date/time: 2026-06-29 · 1:15pm-2:15pm - Track/room: Workshops Day 1 · Track 3 - Speaker(s): Harshul Jain, Tanmay Sah - Session type/status: sponsor · confirmed - Track: Workshops Day 1 - Room: Track 3 - Session type: sponsor - Status: confirmed ## Session Description Most engineers using LLMs can call an API. Far fewer can explain why their model is slow, why it's running out of memory, or how the inference engines powering every major LLM API actually work. This workshop walks through the full inference stack — from how a transformer generates a single token to serving billions of tokens a day with vLLM, SGLang, TensorRT-LLM, Ray, and KServe/llm-d. 60% explanation with live demos, 40% hands-on exercises.
Speaker And Company Context
- Harshul Jain — Senior Software Engineer - ML/AI at Audible.
- Tanmay Sah — Senior Quantitative Modeler at Zions Bancorporation.
Topics Covered
- Topic links are pending transcript-backed classification.
Derived Links And Source Material
Novel Concepts / Clever Methods
- No highlighted novel concept has been detected yet.
Evidence Boundary
This synthesis is based on the official schedule and linked source pages. It should be revisited when exact session recordings or transcript-backed secondary sources are available.