---
title: "Slides: Voice AI: when is the \"Her\" moment? — Neil Zeghidour, CEO, Gradium AI"
category: "slides"
video_id: "P_RI1kCkRbo"
sourceLabels: ["Public YouTube video frames", "Public YouTube metadata"]
---

# Slides: Voice AI: when is the "Her" moment? — Neil Zeghidour, CEO, Gradium AI

## Source Video
[Voice AI: when is the "Her" moment? — Neil Zeghidour, CEO, Gradium AI](https://www.youtube.com/watch?v=P_RI1kCkRbo)

## Relationship To World's Fair 2026
These slides are extracted from a public AI Engineer YouTube video connected to World's Fair 2026. Speaker-matched clips are supporting context unless later confirmed as exact session recordings; official livestream recordings are day-level/event-level source material.

## Related Scheduled Sessions
- No individual scheduled session mapping has been assigned yet; treat this as an event livestream deck.

## Extracted Slides
![[assets/slides/P_RI1kCkRbo/slide-001.jpg]]

OCR text:

> PLATINUM SPONSORS
> Braintrust WorkOS OpenAI

![[assets/slides/P_RI1kCkRbo/slide-002.jpg]]

OCR text:

> aa raat? ; er ; a
> * * oA Tdps
> ‘ . no a a NC) VAsteLnT(G (OU
> ky en ae ee oe :
> a we eC ae CEO & Cofounder
> a a a £ “ Gradium
> ee i ae
> brig ete oar oon
> soos need ae
> ; y
> N
> a Google DeepMind
> cee, tort

![[assets/slides/P_RI1kCkRbo/slide-003.jpg]]

OCR text:

> ; coat ;
> a ra .
> ~ a ae a ; a Pa ot
> Pa 7 5 - - ne i : : eb,
> * i i coe Fi a F mee { *: 7 be en ; :
> at ¥ a ene i, a cae . 2 eS
> Lar aed nr! ‘ & 4 f 4 aes ana * ary
> A By 7 vs xs J . Le
> eNmaeereamraauniock the unrealized potential of voice A a ae a a
> i
> flud natura voce as the new owerface for A ar ; 8 . a !
> (ve tra vonce tode's basically SEE PTS, S28) os ; ; ea . woe a vee s & .
> F es oor
> , =
> rn i
> oe €3 Braintrust €} WorkOS OpenAl
> a tenor

![[assets/slides/P_RI1kCkRbo/slide-004.jpg]]

OCR text:

> From ResearchtoProduction
> Kyutai Breakthrough
> Paris-basedopenscience Al lab. Founded2023withE300M. Moshi:firstspeech-native Research
> AIE Team from MetaFAIR,Google DeepMind,and Inria. Hibiki-Zero:real-time
> speech-to-speech translation.
> moshi.chat by/kyutai Buito bringresearchtoproduction.$70Mraised in2025. Gradium:Scaling Impact Pocket-TTS:CPUmodelforTTS
> AIEngineer
> AIEngineer EUROPE

![[assets/slides/P_RI1kCkRbo/slide-005.jpg]]

OCR text:

> AIE
> AI Engineer
> Engineering the future of AI

![[assets/slides/P_RI1kCkRbo/slide-006.jpg]]

OCR text:

> AIE
> Engineering the future of AI
> AIEngineer

![[assets/slides/P_RI1kCkRbo/slide-007.jpg]]

OCR text:

> In real life...
> AIE
> AI Engineer
> EUROPE

![[assets/slides/P_RI1kCkRbo/slide-008.jpg]]

OCR text:

> In real life...
> ai-PULSE
> AIE
> byScdewoy
> Engineering the future of Al
> AIEngineer
> COROFL

![[assets/slides/P_RI1kCkRbo/slide-009.jpg]]

OCR text:

> Engineering the future of AI

![[assets/slides/P_RI1kCkRbo/slide-010.jpg]]

OCR text:

> Whatismissing?
> Latency
> Contextual Fillers
> User
> AIE
> Agent
> STT
> LLMFillers
> TTS
> Tool Calling
> TTS
> Engineering the future of Al
> AIEngineer

![[assets/slides/P_RI1kCkRbo/slide-011.jpg]]

OCR text:

> page aineaea nena aeresess sewer
> . on ID wamemen: ae @ th @
> an > § nines, \ eres
> a * a _ , aa al are ° . ie
> Fy ry - Zz Leia Pa a
> bd * ae se aan a nae aoa
> lelos eic tk Mae | leis Ear eee ee: ee a ol. 8
> _ €3 Braintrust €) WorkOS OpenAl

![[assets/slides/P_RI1kCkRbo/slide-012.jpg]]

OCR text:

> On, that's 2 great top<! Yeah, I'd hove to help you b
> 
> Peo ‘
> bd bd
> ere
> bd ad
> 
> * ae bd
> 
> “ ‘=> . Ne |
> . eee Wee te te oa! :
> a) | | Al Engineer |
> iD SUL ela
> [nengne_| ;

![[assets/slides/P_RI1kCkRbo/slide-013.jpg]]

OCR text:

> lie glayed:
> AIE
> /.kyutai
> Engineering the future of Al
> AIEngineer

![[assets/slides/P_RI1kCkRbo/slide-014.jpg]]

OCR text:

> Moshi Lessons Learned
> What Works Well
> Full-Duplex Conversation: Natural, uninterrupted
> dialogue flow
> Real-time Interaction: Truly conversational AI, user and
> AI speak simultaneously
> Challenges
> A research prototype, not a real agent
> Lack of observability and reliability
> Does not convey or understand empathy well
> Engineering the future of AI

![[assets/slides/P_RI1kCkRbo/slide-015.jpg]]

OCR text:

> GradiumPhonon Real-Time InferenceonCPU
> CPUInference Weights WER SpeakerSim
> Faster than real-time with no perceptible delay Phonon ~100M 1.48% 56.37%
> AIE Kani TTS 2 450M 4.97% 40.73%
> Personalization NeuTTS Air 552M 2.18% 47.51%
> Multilingual+canreproduce any voicefroma short referenceclip,noretrainingrequired Kokoro NeuTTSNano 229M 100M 1.71% 0.90% 40.15%
> SOTAperformance
> Benchmarkon seed-tts→
> Engineering the future of Al
> AIEngineer

![[assets/slides/P_RI1kCkRbo/slide-016.jpg]]

OCR text:

> Cn “Eye +h O® Be Fak wweun
> ED eo rete mene se ws
> * Te ee ree entree sated Le QP ESE ORE MUNN enn ad ERE abe a a6 @.% @
> JINR RT coca Ore Bo Hn ee 6 Oo
> Eo oF" 4.59 SAGE IR OAS G 1 ewe ceRE ee oro wee Cgcetginite =
> »
> z=
> * bg *”
> ace =
> Ls * * “hepath fopaatdfan yc ce espe y fort oped aerators bes a 2
> loan ial . , —
> Bi science and engineering Scene
> F ; >
> ca [od ee CCS SE Oca SOMES A LO OOO BO SO LTRS MESS
> -_ - *
> Taner)
> = sere rte teen
> wBagCF®eO<Sse BBO. - "474° OB * Sse Le
> ‘a . Engineering the fut f Al
> I - Hl

![[assets/slides/P_RI1kCkRbo/slide-017.jpg]]

OCR text:

> AI Engineer
> EUROPE
> HTTPS://AI.ENGINEER

## Slide-Derived Subjects To Review
Subject extraction uses video title, related session titles/descriptions, transcript context, and OCR text when available. OCR is best-effort and should be reviewed against the embedded slide images.
