---
title: "Dense Slides: Training Agentic Reasoners — Will Brown, Prime Intellect"
category: "slides"
video_id: "PbHm2qKnu10"
sourceLabels: ["Captured video frames", "Local OpenCV slide-region detection"]
---

# Dense Slides: Training Agentic Reasoners — Will Brown, Prime Intellect

## Source Video
[Training Agentic Reasoners — Will Brown, Prime Intellect](https://www.youtube.com/watch?v=PbHm2qKnu10)

## Method
This deck is slide-only. The existing captured video frame set supplies candidate frames, then local OpenCV rejects sponsor/title/speaker-only frames, crops visible slide surfaces, deduplicates, and saves the cropped slide images.

## Cropped Visible Slides
![[assets/dense-slides/PbHm2qKnu10/slide-001.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/dense/PbHm2qKnu10/slide-001.html)
- AI slide classifier: `content_slide` confidence `0.99`
- Text source: advanced OCR `rapidocr-live/border-trim/opencv-adaptive`.
- OCR decision: ready — Content slide with chart text and an embedded UI screenshot containing dense small text.

Slide text:

> RL kinda works now
> AIE Q DNeDSk.Al-Zero AbE aCCuraCy during trlnlng Hlrtet Summury >NvDu Corp 118.58 uso -20.64 (-14.83%) + past 5 dny
> Caha 2? Lr. s or Pu Gfs + Dechtr A+ Noun +20 (4 +2 c0 1+ 74)
> 10 50 11 64t YTD 5Y LL
> 150
> 141*0-2160 14 143++d 4++-11 rl-r*+(o*+1+ 120 01
> $Hro 000* 1000 Ito 24n
> Microsoft smop

![[assets/dense-slides/PbHm2qKnu10/slide-002.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/dense/PbHm2qKnu10/slide-002.html)
- AI slide classifier: `content_slide` confidence `0.99`
- Text source: advanced OCR `rapidocr-live/bright-screen/opencv-adaptive`.
- OCR decision: ready — Content slide with a paragraph of small body text and a video thumbnail screenshot containing readable text.

Slide text:

> the big labs are all doing it
> AIE learning computo = bettor porformance' trond observod in GPT-serics pretraining. By retracing the scaling path--thia time in RL-we've pushed an additional order of magnitude in both tralning performance keeps chmbing. Continuing to scale reinforcement Throughout tha deveiopment of OpenAl o3, we've observed that large-scale reinforcement learning oxhibits the sama "more computa and inferonce-tirmo roasoning. yet still seo cloar performance gains, valldating that the models' porformance continues to lmprove the more they're allowed to think At equa! in ChatGPT-and wo'vo validatod that if wo lot k think longer, its Latcncy and cost with OpenAl ol, o3 delirvers hlgher performance Is Rl + Llms enough for AGl? - 123K views · 13 days ag0 Sholto Douglas & Trenton Bricken AND NEXT WHAT COMES 2:24:02
> Microsoft smop

![[assets/dense-slides/PbHm2qKnu10/slide-003.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/dense/PbHm2qKnu10/slide-003.html)
- AI slide classifier: `content_slide` confidence `0.99`
- Text source: advanced OCR `rapidocr-live/full`.
- OCR decision: ready — Content slide with multiple dense text regions, code screenshots, and a diagram that are better handled by OCR.

Slide text:

> agents are also a thing TaskExecutionProcess MANUSACAGENT
> Dlenaare
> pupnd
> AIE alcoes ts Clode Code researchsrevieei ela'fer bes fer getting stere 25003000.By sectingthsrgont efntredehen evalute the plte and se itirs cearer o idey wthapprmate cordiex1500 to2000and Let'smanually crop and inspecti
> Analyzing data
> slt.aist`gre)
> -0.5.899.5. 799.5
> Analyzing image
> Microsoft smol?

![[assets/dense-slides/PbHm2qKnu10/slide-004.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/dense/PbHm2qKnu10/slide-004.html)
- AI slide classifier: `content_slide` confidence `0.98`
- Text source: agent_vision.

Slide text:

> they’re kinda the same thing actually

![[assets/dense-slides/PbHm2qKnu10/slide-005.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/dense/PbHm2qKnu10/slide-005.html)
- AI slide classifier: `content_slide` confidence `0.99`
- Text source: advanced OCR `rapidocr-live/full`.
- OCR decision: ready — Content slide with code, chart, file list, and a timeline diagram; OCR is appropriate.

Slide text:

> towardseverythingasync
> AIE returnaaittode_asyncio,gathert rollowt_tasks-[ self._rn_single（sephere,client,model,pronpt,aniver,sapling_args,args） fetpronpt,answer-inzippronpts,ansers) *rollout_tasks, total-len(prompts), descof'Running(len(pronptsl)rollouts async_batch_generator.py async_dataloader_wrapper.py
> grpo_config.py
> grpo_trainer.py
> GPU1 GPU2 GPU3 GPU4 1-8 9-16 17-24
> Time
> (up to four steps), Prime-RL matches the performance of synchronous baselines. us DeepScaleR trainingvs asynchronous Prime-RL under varyingasynchrony levels.
> aws

![[assets/dense-slides/PbHm2qKnu10/slide-006.jpg]]

- Recreated text/layout view: [open HTML recreation](/assets/slide-recreations/dense/PbHm2qKnu10/slide-006.html)
- AI slide classifier: `content_slide` confidence `0.99`
- Text source: advanced OCR `rapidocr-live/bright-screen/contrast`.
- OCR decision: ready — Content slide with code on the left and a product UI screenshot on the right; OCR will likely outperform manual transcription.

Slide text:

> SFT warmup + small models = fun on just a few GPUs
> PRumeintllet CreatenewGPU Cluster
> AIE results=vf_env.evaluate( client=client, model=model_nane, sompling_args=sonpling_args, 200 141 Ga DO N100 B0 C Ala 00NS
> num_saeples=nun_sanples 12.14
> hub_odel_id"0wen2.5-7B-Math-Python-SFT", A100 0028
> 公
> PCxSOM 1MS
> trainer=SfTrainer( nodel=nodel,
> argseargs,
> train_dataset=datasettype:icnore
> GH200 A100
> trainer.train() 6) (/
> 入7, 344
> Microsoft smol ai


Classification audit: `raw/sources/slide-ai-classification/dense/PbHm2qKnu10/audit.json`
