---
title: "Slides: Training Agentic Reasoners — Will Brown, Prime Intellect"
category: "slides"
video_id: "PbHm2qKnu10"
sourceLabels: ["Public YouTube video frames", "Public YouTube metadata"]
---

# Slides: Training Agentic Reasoners — Will Brown, Prime Intellect

## Source Video
[Training Agentic Reasoners — Will Brown, Prime Intellect](https://www.youtube.com/watch?v=PbHm2qKnu10)

## Relationship To World's Fair 2026
These slides are extracted from a public AI Engineer YouTube video connected to World's Fair 2026. Speaker-matched clips are supporting context unless later confirmed as exact session recordings; official livestream recordings are day-level/event-level source material.

## Related Scheduled Sessions
- No individual scheduled session mapping has been assigned yet; treat this as an event livestream deck.

## Extracted Slides
![[assets/slides/PbHm2qKnu10/slide-001.jpg]]

OCR text:

> aWS
> 
> eee)
> @®Graphite WW Windsurf 4 MoneobB
> Mdaily £3 augment code WorkOS

![[assets/slides/PbHm2qKnu10/slide-002.jpg]]

OCR text:

> training agentic reasoners
> will brown
> @willecbb
> research lead @ prime intellect
> pr world’s fair 2025
> ‘i a
> , | ; a Microsoft ary?

![[assets/slides/PbHm2qKnu10/slide-003.jpg]]

OCR text:

> the big labs are all doing it
> Continuing to scale reinforcement
> learning AND - #
> v7) |
> Throughout the development of OpenAl 03, we've observed that
> large-scale rewnforcement learning exhityts the same “more oF @) | = Ss
> compute © better performance™ trend observed in GPT-series
> pretrainng. By retracing the scaling path—ths time in RL-we've N = xX a ‘ ¥ x
> pushed an additional order of magnitude in both training 9:24:02
> compute and inference-time reasoning, yet still see clear ¢" ‘nn
> performance gams. validating that the models’ performance
> continues to improve the more they're allowed to think. At equal Is RL +LLMs enough for AGI? - :
> latency and cost with OpenAl o1, 03 delrvers higher performance .
> in ChatGPT— and we've validated that if we fet it think longer, its Sholto Douglas & Trenton Bricken
> Performance keeps chmbing. 123K views + 13 days ago
> FY
> t4
> - , rai)
> — PLES ION

![[assets/slides/PbHm2qKnu10/slide-004.jpg]]

OCR text:

> agents are also a thing TaskExecutionProcess INYISNY
> Iwanttocropamal bounding boaround the lcene
> AIE atertedi pp Let's monually crop and inspecti.
> OLAC Analyzing data
> 22035)
> slt.adet'efr)
> （-0.5,899.5,799.5.-0.5）
> eu uz/euy
> Microsoft smol?

![[assets/slides/PbHm2qKnu10/slide-005.jpg]]

OCR text:

> they're kinda the same thing actually
> AIE LM Ca
> Corporate between thispictuteand thispicture. needsy outofind the differences
> aws

![[assets/slides/PbHm2qKnu10/slide-006.jpg]]

OCR text:

> .
> SFT warmup + small models = fun on just a few GPUs
> as ‘ St ee |
> results = vf_env.evaluatet S cares)
> clientectient, ,
> eodel=aodel_nane, He
> sanpting_args:saapling_args,
> Nus_samples=num_sanptes . i
> ) ,
> hub _sodel_id="Qwen?. $-78-Math-Python-SFT™, _ -
> }
> trainer > SFTTrainert ;
> mode l=nodel, : :
> ergsrargs,
> train dateset=dataset @ type: ignore Py
> ) . r
> trainer. trarn()
> “ aws
> | a
> x
> wf me

## Slide-Derived Subjects To Review
Subject extraction uses video title, related session titles/descriptions, transcript context, and OCR text when available. OCR is best-effort and should be reviewed against the embedded slide images.
