---
title: "Slides: How to build world-class AI products — Sarah Sachs (AI lead @ Notion) &  Carlos Esteban (Braintrust)"
category: "slides"
video_id: "6YdPI9YbjbI"
sourceLabels: ["Public YouTube video frames", "Public YouTube metadata"]
---

# Slides: How to build world-class AI products — Sarah Sachs (AI lead @ Notion) &  Carlos Esteban (Braintrust)

## Source Video
[How to build world-class AI products — Sarah Sachs (AI lead @ Notion) &  Carlos Esteban (Braintrust)](https://www.youtube.com/watch?v=6YdPI9YbjbI)

## Relationship To World's Fair 2026
These slides are extracted from a public AI Engineer YouTube video connected to World's Fair 2026. Speaker-matched clips are supporting context unless later confirmed as exact session recordings; official livestream recordings are day-level/event-level source material.

## Related Scheduled Sessions
- No individual scheduled session mapping has been assigned yet; treat this as an event livestream deck.

## Extracted Slides
![[assets/slides/6YdPI9YbjbI/slide-001.jpg]]

OCR text:

> INNOVATIONPARTNER
> aws
> PLATINUMSPONSORS
> Graphite
> WWindsurf
> MongoDB
> daily
> augment code
> WorkOs

![[assets/slides/6YdPI9YbjbI/slide-002.jpg]]

OCR text:

> Tiel!
> ‘ile Te Tal
> a
> He | 3
> mt! oe:
> “r , (eta
> eu . ¥
> ~ a - ;
> a
> P . 7 a
> | re —) ee re 2
> eal a a. ae cae
> a ie i i
> 7 -
> a i \ , P 7 = -
> mnie; | 1
> ian! a

![[assets/slides/6YdPI9YbjbI/slide-003.jpg]]

OCR text:

> co r
> Sign up for Access Join the
> SieciIAIGRersi workshop eT
> materials channel
> Ope sa Ease Ena
> eae eet aioe ey
> fates (alpeee at lor
> _ fr}
> | Eel eM r Ce
> 1 7

![[assets/slides/6YdPI9YbjbI/slide-004.jpg]]

OCR text:

> BRAINTRUST,DEV
> AIE
> HowNotion develops
> world-class Alfeatures
> CONFIDENTLAI
> Microsoft
> smol?
> World'sFair

![[assets/slides/6YdPI9YbjbI/slide-005.jpg]]

OCR text:

> eee Alarms 2 Doe 6h ON ewe Oot oe
> e §=Connected workspace Se)
> ner
> for documents, >
> ; ee ve Fe
> xnowledge bases, prarect ”
> menedement ’
> ~J Docs
> en OOM ral ore Ramenst ares Sa Ses (Birmrenrases = co . Eg
> globally, from. startups to B:rees monae anoourero row a
> S ¢ Vacation policy a
> raise OtaRS Sood @ Ortadare relntion properten "
> 1, Recretng new support rept a
> 3S tow Acme uses Stripe a
> Can @ poh Rin Col ance Xonzs lero raz] Taga Wtceme’ ao
> loreichee] Berm sr co rssa Lexa lets) ‘Witonseing Wienevs nies a
> of Report monthyy anatytes. =
> of Creanng help center docs ° >
> » aa ~
> i :
> / eased

![[assets/slides/6YdPI9YbjbI/slide-006.jpg]]

OCR text:

> BRAINTRUST.DEV
> Alforwork
> AIE
> Newdashboard launch
> Project Overview
> Key Objectives
> Microsoft
> smol?
> World'sFair

![[assets/slides/6YdPI9YbjbI/slide-007.jpg]]

OCR text:

> z ®
> Early leaders in Generative Al
> Nov 2022 Jun 2023 Nov 2023
> Vai ccle Al Autofill: prompt AL Q&A: RAG
> generate & edit * across hundreds * across your Notion
> relolensy fol ore]e lee ale workspace
> oo aws

![[assets/slides/6YdPI9YbjbI/slide-008.jpg]]

OCR text:

> Evaluation challenges
> Iniual workflow:
> e Large JSONL datasets difficult to manage and version
> AE e Human evaluators costly and overwhelmed by volume
> > Need for scalable and efficient evaluation
> fe »
> » a
> | , 24]
> oe ma Microsoft §GOU
> ~
> [ World's Fa

![[assets/slides/6YdPI9YbjbI/slide-009.jpg]]

OCR text:

> With Braintrust
> 1. Decide on an improvement
> 2. Curate targeted datasets (logs + handcrafted examples)
> 3. Tie datasets to specific scoring functions (heuristics, LLM-as-judge, human review)
> 4. Run evals, inspect details results
> 5. Iterate until ready to ship
> cd
> A fm™
> i ae
> ho aws
> Pe iz, all

![[assets/slides/6YdPI9YbjbI/slide-010.jpg]]

OCR text:

> Fast feedback for fast development
> Modular stack enables rapid iteration and
> continuous evaluation
> ; a SPECIALISTS
> LLM-as-a-judge system run by Al data specialists
> e Design custom evaluation criteria LLM-AS-A-JUDGE
> e Analyze real user behaviors, improving i)
> prompts beyond benchmarks
> — ; ARON a
> e Evaluate and deploy new models (OpenAl,
> Anthropic, Google, open-source)
> DEPLOY T
> Continuous evaluations ensure quality, catch
> early, validate improvements
> i
> 4 aws
> ae 3 a

![[assets/slides/6YdPI9YbjbI/slide-011.jpg]]

OCR text:

> BRAINTRUST.DEV
> Outcomes
> AIE
> 10x 50%
> More issues triaged and fixed/day Increase in Al product quality
> OOAFICENTLAL
> aws
> Worid

![[assets/slides/6YdPI9YbjbI/slide-012.jpg]]

OCR text:

> BRAINTRUST.DEV
> AIE
> Q&A
> CONFIDENTIAL

![[assets/slides/6YdPI9YbjbI/slide-013.jpg]]

OCR text:

> ry Tasks (Promots, Batra Messages, Aqaits + Toei?
> Cee Ore Seto
> Cy Scares Agtorsals OOM adiges Coe pudaes
> e «(Off ne BOn re bvas
> . Playground ys Peper mats
> 
> {Intro to Evalsin UL Actuity]
> oy Sots ae Soe OS |
> e deploy Unreerased ALApA (2 mi nutes)
> 
> Liner om com er nS] 01 Geer Cok hated
> ° (ene pea ra a no
> ° Oeste Catal Chaco le Marr aea a ETL LES
> 
> {Logaqing & Online Scormnyg » Activity)
> O LATOR CordL er aod Ob toa ted hss | Ove oC nae OC ta)
> 
> ome a q& Human Reviesy - Actuty]
> 
> = ne SS

![[assets/slides/6YdPI9YbjbI/slide-014.jpg]]

OCR text:

> AO) ales) (c 101.6
> How do you currently
> evaluate your Al systems?
> |
> : oJ 7 a Microsoft ary?

![[assets/slides/6YdPI9YbjbI/slide-015.jpg]]

OCR text:

> a Ps
> Why evals? Evals help answer questions
> Model selection How dees pa Cost efficiency
> "Which LLM is the erform across "Can we achieve high
> best choice for our dares real-world performance without
> needs?" scenarios?" excessive costs?"
> . Debugging & regression
> . Booey et cy Feedback integration detection
> ' des the Al reliably “Are we learning from “How do we know
> reflect our company's users and improving when something
> voice and | iteratively?” breaks or gets
> standards? worse?"
> = V1 rs)
> _ a. a Microsoft =U

![[assets/slides/6YdPI9YbjbI/slide-016.jpg]]

OCR text:

> rd
> How evals can help your business
> Orel dev time Rapid iteration cycies & local testing an multiple LLMs seamlessly
> Sqs1o1ULer> (ore kci as) Automated evais replace manual review alowing faster tteration / release
> 
> E alate alele q Ua na, Real time monitoring & compliance to reduce risk and improve Cx
> Sca le team S Enabie non-technical collaboration to build tne best Al elon
> 
> Ive never seen a workfiow transformation like the one that incorporates
> 
> evals into ‘mainstream engineering’ processes before. It’s astonishing.
> 
> fone
> ¥
> “a
> | _—_ 7

![[assets/slides/6YdPI9YbjbI/slide-017.jpg]]

OCR text:

> ce, e
> 
> What is an Eval?
> Definition: An Eval (short for evaluation) is a
> 
> structured test that checks how well your Al
> 
> system performs. It helps you measure
> quality, reliability, and correctness across
> scenarios.
> 
> M1 :
> 
> - a oy a Microsoft ary?

![[assets/slides/6YdPI9YbjbI/slide-018.jpg]]

OCR text:

> 3 Ingredients in an Eval
> DATASET SCORER
> ~The logie behind your
> The code or prompt you evals.
> want to evaluate. ~
> Can be an
> -It can ranqe tron a LLM as a judge or full
> single prompt to an A set of real-world Cede function
> rahe a mCMmeTe Srp GME Ola Ge a CrvaN examples, 7 ; ;
> The scorer will give a
> PORTO P Reem Mes TLCL aeT “These are your test score of B-108 to each
> an output. cases. row in the dataset.
> — ‘ mE okor to ES O00 Ue
> oo 7

![[assets/slides/6YdPI9YbjbI/slide-019.jpg]]

OCR text:

> Offline Evais
> Sanat | ns
> e = Structures testing of Al mode's us na e Realtree tacing aad morstonng of apoir
> predefined datesels. Pete
> PSna Teta Te Mia ters fedora a old iih I Mal en eT| a Tors iata a] © Logs of mecel inputs, oLtots, and interred. ate
> Ce Daren Tera iciees
> ar evan eed
> Wes a
> e@ Proactive ident fy and ceso ve issues Crear: [elsiol ttm Sice)s ict ES ME GaLO Oth exis BI et Le] G1 0Ts ker toa]
> before deploy rent Fars] ps C0] Rtas Cacres@ne od S76] Ors Loc Sateen? MGT
> rerOn How
> ©) Create and marace evaliation datasets ©) AGtomaticdaby (ace cabs usiag proxy or SOK.
> PEE BIO HITCRITCT ETC KSiT aes e ose fers and ves p Granteust UE to araly re
> CS TReMaraaa ade] MOOSE OCA TEETIa Tet aTemare iTS Tae) | eens ;
> a Tete ee!| e §6Carvertiogs directly into datasets for offlre eva s
> a x Wesel) aS 6000 Oe
> ‘ —

![[assets/slides/6YdPI9YbjbI/slide-020.jpg]]

OCR text:

> BRAINTRUST,DEV
> What should I improve?
> 7
> AIE Goodoutput High score Improveevals Lowscore
> ndinopeg Improveevals ImproveAlapp
> WorkfsFa aws

![[assets/slides/6YdPI9YbjbI/slide-021.jpg]]

OCR text:

> Fear ys Peay Sree CO UROET Ta TORS gy RO Co Se e. Loe Le . °
> PoC to pirate pepo se
> i arery
> 
> Dee ra es ocean ia oa Oy a —
> ed cee ie ae a ar
> fi or ay eae arn er mac eer Cine bo ee
> ar a ee ae ns le were
> Coesah te ee ee
> 
> ° Sees
> 
> . ce aves ER
> 
> Pe On oer tee ;
> (ed we mt as ee Oa or (once med Lea er
> hres Perea ey | ee Coes eee
> Tee esas eae ec Re en
> CCSa Ens Re Oare NS [ESM SCE cS
> Partai
> 
> ~ —,

![[assets/slides/6YdPI9YbjbI/slide-022.jpg]]

OCR text:

> a , ®
> Datasets - Tips
> 
> Cero) ©: 18 ecvn irs] mene | Con
> focus on building a feedback loop rather than
> a perfect dataset
> 
> e = =Never Stop Iterating:
> Use Logs to capture more edge cases and
> create more holistic Evals
> 
> e Implement Human Reviews:
> use human review to establish ground truth
> Especially when using expected column
> 
> ° 4
> loro

![[assets/slides/6YdPI9YbjbI/slide-023.jpg]]

OCR text:

> BRAINTRUST.DEV
> Scorer Types
> AIE
> Code-based scorers
> - Exact or binary conditions
> - Numeric comparisons
> - Structured or factual checks
> LLM-as-a-judge scorers
> - Subjective or contextual feedback
> - Human-like interpretation
> - Improvement across multiple drafts

![[assets/slides/6YdPI9YbjbI/slide-024.jpg]]

OCR text:

> Scorer Tips
> Use a higher-quality model for scoring, even if the prompt uses a cheaper
> model. Scorers benefit from better reasoning and nuance.
> Treat scorers like judges: evaluate intent match, style accuracy, and
> overall output quality: -not just correctness.
> Break scoring into multiple focused scorers (e.g., accuracy, Creativity,
> formatting) to pinpoint issues.
> Test scorer prompts in the Playground before use. Try strong and weak
> outputs to refine scoring reliability.
> Avoid overloading the scorer prompt with context. Focus it on the relevant
> input and output for fair, consistent evaluation.
> [wire |
> ce

## Slide-Derived Subjects To Review
Subject extraction uses video title, related session titles/descriptions, transcript context, and OCR text when available. OCR is best-effort and should be reviewed against the embedded slide images.
