Markdown source

Build realtime multimodal agents with Gemini Live

Conference Context

Session Description

The Gemini Live API is incredible versatile when it comes to building realtime AI experiences. From live translation across 2000 different language pairs to building realtime multimodal agents that can work across text, audio, and vision. This workshop gets you from zero to fully conversational agent in a matter of hours.

Media Evidence

From Transcription to Live Music: Gemini's Audio Stack — Thor Schaeff, Google DeepMind (speaker-match related prior/adjacent AI Engineer video; captions: English auto-captions).

These are phone-photo slide captures from the Google Photos AIE Slides album. They are supporting slide evidence and do not override official schedule fields.

Evidence Graph

This evidence graph is generated from currently linked source material: official schedule text, related video pages, cached transcripts, visible slide text, dense/reconstructed slide pages, and AI slide-classification audits.

Media Signals

Agent Reading Notes

Use these signals to refine the synopsis, topic links, people/company context, and method notes. If a source is a related external video rather than an exact official recording, keep it framed as supporting evidence.

Transcript Status

Related video transcript availability: English auto-captions. Treat this as supporting context, not a recording of this exact scheduled session unless later confirmed. Not fetched yet.

People

Supporting Slides

Slide Evidence

Synthesis

Synthesized Breakdown

Build realtime multimodal agents with Gemini Live ## Conference Context - Date/time: 2026-06-30 · 10:45am-11:05am - Track/room: Workshops Day 2 · Track 4 - Speaker(s): Thor 雷神 Schaeff - Session type/status: session · confirmed - Track: Workshops Day 2 - Room: Track 4 - Session type: session - Status: confirmed ## Session Description The Gemini Live API is incredible versatile when it comes to building realtime AI experiences. From live translation across 2000 different language pairs to building realtime multimodal agents that can work across text, audio, and vision. This workshop gets you from zero to fully conversational agent in a matter of hours. ## Media Evidence From Transcription to Live Music: Gemini's Audio Stack — Thor Schaeff, Google DeepMind (speaker-match related prior/adjacent AI Engineer video; captions: English auto-captions).

Speaker And Company Context

Topics Covered

Derived Links And Source Material

Novel Concepts / Clever Methods

Evidence Boundary

This synthesis is based on the official schedule and linked source pages. It should be revisited when exact session recordings or transcript-backed secondary sources are available.