---
title: "How to Connect AI to Billions of Legal Documents"
category: "talks"
date: "2026-06-29"
time: "2:25pm-2:45pm"
track: "Search & Retrieval"
room: "Track 3"
speakers: ["Simon Eskildsen", "Jacob Lauritzen"]
sourceLabels: ["Official conference schedule", "Public YouTube metadata"]
scheduleTrack: "Search & Retrieval"
scheduleRoom: "Track 3"
scheduleLabels: ["Search & Retrieval", "Track 3", "session", "confirmed"]
---
# How to Connect AI to Billions of Legal Documents

## Conference Context
- Date/time: 2026-06-29 · 2:25pm-2:45pm
- Track/room: Search & Retrieval · Track 3
- Speaker(s): Simon Eskildsen, Jacob Lauritzen
- Session type/status: session · confirmed

- Track: Search & Retrieval
- Room: Track 3
- Session type: session
- Status: confirmed

## Session Description
Legora’s foundational engineering challenge is connecting frontier LLMs to billions of legal documents so the models can efficiently solve end-to-end legal workflows without burning extra tokens. We’ll share the retrieval architecture we built with turbopuffer that achieves: 1. Strict data isolation across millions of legal cases in a very security-conscious domain 2. Predictable search performance (<100ms p90 latency) on large contexts 3. High retrieval quality (95%+ recall@10) with fewer agent loops We’ll retrospect on two architectures that failed to achieve all 3 (and why), and the key design factors that make the current solution work at our scale. Practical takeaways include: - How to evaluate per-tenant vs shared-index retrieval under strict data isolation - How to efficiently index and retrieve context to maximize relevance per input token - How to build a highly intelligent AI application when your inference budget is constrained

## Media Evidence
[Agents need more than a chat - Jacob Lauritzen, CTO Legora](https://www.youtube.com/watch?v=XNtkiQJ49Ps) (speaker-match related prior/adjacent AI Engineer video; captions: English auto-captions).

- Source video: `youtube-XNtkiQJ49Ps`
- Slide deck: [[youtube-XNtkiQJ49Ps-dense-slides|Dense Slides: Agents need more than a chat - Jacob Lauritzen, CTO Legora]] — 7 visible slide image(s); 7 HTML recreation(s).
![[assets/dense-slides/XNtkiQJ49Ps/slide-001.jpg]]
![[assets/dense-slides/XNtkiQJ49Ps/slide-002.jpg]]
![[assets/dense-slides/XNtkiQJ49Ps/slide-003.jpg]]
- Additional slide evidence: [[youtube-XNtkiQJ49Ps-slides|Slides: Agents need more than a chat - Jacob Lauritzen, CTO Legora]], [[youtube-XNtkiQJ49Ps-reconstructed-slides|Reconstructed Slides: Agents need more than a chat - Jacob Lauritzen, CTO Legora]]
- Slide-derived themes for `youtube-XNtkiQJ49Ps`: human, than, chat, jacob, collaborative, legal, professionals, customers.

## Evidence Graph
This evidence graph is generated from currently linked source material: official schedule text, related video pages, cached transcripts, visible slide text, dense/reconstructed slide pages, and AI slide-classification audits.

### Media Signals
- `youtube-XNtkiQJ49Ps` — 9 slide-derived text signals
- Slide-derived themes for `youtube-XNtkiQJ49Ps`: human, than, chat, jacob, collaborative, legal, professionals, customers.
- Evidence links for `youtube-XNtkiQJ49Ps`: [[youtube-XNtkiQJ49Ps]], [[youtube-XNtkiQJ49Ps-slides]], [[youtube-XNtkiQJ49Ps-dense-slides]], [[youtube-XNtkiQJ49Ps-reconstructed-slides]]

### Agent Reading Notes
Use these signals to refine the synopsis, topic links, people/company context, and method notes. If a source is a related external video rather than an exact official recording, keep it framed as supporting evidence.

## Transcript Status
Related video transcript availability: English auto-captions. Treat this as supporting context, not a recording of this exact scheduled session unless later confirmed. Not fetched yet.

## People
- [[simon-eskildsen]]
- [[jacob-lauritzen]]

## Supporting Slides
- [[youtube-XNtkiQJ49Ps-slides]] — extracted from the related public AI Engineer video.

## Slide Evidence
- Slide-only cropped deck: [[youtube-XNtkiQJ49Ps-dense-slides]] (7 viable slide images).
- Related slide/OCR pages:
- [[youtube-XNtkiQJ49Ps-dense-slides]]
- [[youtube-XNtkiQJ49Ps-reconstructed-slides]]
- [[youtube-XNtkiQJ49Ps-slides]]
- Slide-derived terms: `legora`, `chat`, `than`, `company`, `jacob`, `lauritzen`, `trust`, `searching`, `reading`, `braintrust`, `workos`, `openal`, `files`, `file`, `humans`, `alengineer`, `vecoea`, `collard`

## Synthesis
### Synthesized Breakdown
# How to Connect AI to Billions of Legal Documents ## Conference Context - Date/time: 2026-06-29 · 2:25pm-2:45pm - Track/room: Search & Retrieval · Track 3 - Speaker(s): Simon Eskildsen, Jacob Lauritzen - Session type/status: session · confirmed - Track: Search & Retrieval - Room: Track 3 - Session type: session - Status: confirmed ## Session Description Legora’s foundational engineering challenge is connecting frontier LLMs to billions of legal documents so the models can efficiently solve end-to-end legal workflows without burning extra tokens. We’ll share the retrieval architecture we built with turbopuffer that achieves: 1. Strict data isolation across millions of legal cases in a very security-conscious domain 2. Predictable search performance (<100ms p90 latency) on large contexts 3.

### Speaker And Company Context
- [[simon-eskildsen|Simon Eskildsen]] — CEO and co-founder at [[turbopuffer|turbopuffer]].
- [[jacob-lauritzen|Jacob Lauritzen]] — CTO at [[legora|Legora]].

### Topics Covered
- [[agent-security]]
- [[agentic-search]]

### Derived Links And Source Material
- [[youtube-XNtkiQJ49Ps]] — related YouTube source page.
- [[youtube-XNtkiQJ49Ps-slides]] — slide evidence.
- [[youtube-XNtkiQJ49Ps-reconstructed-slides]] — slide evidence.
- [[youtube-XNtkiQJ49Ps-dense-slides]] — slide evidence.

### Novel Concepts / Clever Methods
- No highlighted novel concept has been detected yet.

### Evidence Boundary
This synthesis is based on the official schedule and linked source pages. It should be revisited when exact session recordings or transcript-backed secondary sources are available.
