---
title: "Reinforcement Learning without Verifiable Rewards"
category: "talks"
date: "2026-06-30"
time: "1:30pm-1:50pm"
track: "Posttraining & Midtraining"
room: "Track 9"
speakers: ["Will Brown"]
sourceLabels: ["Official conference schedule", "Public YouTube metadata"]
scheduleTrack: "Posttraining & Midtraining"
scheduleRoom: "Track 9"
scheduleLabels: ["Posttraining & Midtraining", "Track 9", "session", "confirmed"]
---
# Reinforcement Learning without Verifiable Rewards

## Conference Context
- Date/time: 2026-06-30 · 1:30pm-1:50pm
- Track/room: Posttraining & Midtraining · Track 9
- Speaker(s): Will Brown
- Session type/status: session · confirmed

- Track: Posttraining & Midtraining
- Room: Track 9
- Session type: session
- Status: confirmed

## Session Description
Verifiable rewards are the gold standard for RL training, but real-world agent tasks frequently lack clean deterministic evaluation objectives. This talk surveys our efforts to scale RL in non-verifiable settings -- including task synthesis, unsupervised environment design, and automatic judge calibration -- to ultimately enable self-improvement in production, grounded in real-world agent traces and domain-specific context.

## Media Evidence
[Reinforcement Learning for Agents - Will Brown, ML Researcher at Morgan Stanley](https://www.youtube.com/watch?v=JIsgyk0Paic) (speaker-match related prior/adjacent AI Engineer video; captions: English auto-captions).

- Source video: `youtube-JIsgyk0Paic`
- Slide deck: [[youtube-JIsgyk0Paic-dense-slides|Dense Slides: Reinforcement Learning for Agents - Will Brown, ML Researcher at Morgan Stanley]] — 11 visible slide image(s); 11 HTML recreation(s).
![[assets/dense-slides/JIsgyk0Paic/slide-001.jpg]]
![[assets/dense-slides/JIsgyk0Paic/slide-002.jpg]]
![[assets/dense-slides/JIsgyk0Paic/slide-003.jpg]]
- Additional slide evidence: [[youtube-JIsgyk0Paic-slides|Slides: Reinforcement Learning for Agents - Will Brown, ML Researcher at Morgan Stanley]], [[youtube-JIsgyk0Paic-reconstructed-slides|Reconstructed Slides: Reinforcement Learning for Agents - Will Brown, ML Researcher at Morgan Stanley]]
- Slide-derived themes for `youtube-JIsgyk0Paic`: many, pipelines, feedback, best, practices, level, systems, take.

## Evidence Graph
This evidence graph is generated from currently linked source material: official schedule text, related video pages, cached transcripts, visible slide text, dense/reconstructed slide pages, and AI slide-classification audits.

### Media Signals
- `youtube-JIsgyk0Paic` — 8 slide-derived text signals
- Slide-derived themes for `youtube-JIsgyk0Paic`: many, pipelines, feedback, best, practices, level, systems, take.
- Evidence links for `youtube-JIsgyk0Paic`: [[youtube-JIsgyk0Paic]], [[youtube-JIsgyk0Paic-slides]], [[youtube-JIsgyk0Paic-dense-slides]], [[youtube-JIsgyk0Paic-reconstructed-slides]]

### Agent Reading Notes
Use these signals to refine the synopsis, topic links, people/company context, and method notes. If a source is a related external video rather than an exact official recording, keep it framed as supporting evidence.

## Transcript Status
Related video transcript availability: English auto-captions. Treat this as supporting context, not a recording of this exact scheduled session unless later confirmed. Not fetched yet.

## People
- [[will-brown]]

## Supporting Slides
- [[youtube-JIsgyk0Paic-slides]] — extracted from the related public AI Engineer video.

## Slide Evidence
- Slide-only cropped deck: [[youtube-JIsgyk0Paic-dense-slides]] (11 viable slide images).
- Related slide/OCR pages:
- [[youtube-JIsgyk0Paic-dense-slides]]
- [[youtube-JIsgyk0Paic-reconstructed-slides]]
- [[youtube-JIsgyk0Paic-slides]]
- Slide-derived terms: `engineering`, `level`, `reasoning`, `prompt`, `completions`, `models`, `responses`, `completion`, `count`, `better`, `deepseek`, `works`, `rewards`, `next`, `llms`, `chatbots`, `work`, `ai.engineer`

## Synthesis
### Synthesized Breakdown
# Reinforcement Learning without Verifiable Rewards ## Conference Context - Date/time: 2026-06-30 · 1:30pm-1:50pm - Track/room: Posttraining & Midtraining · Track 9 - Speaker(s): Will Brown - Session type/status: session · confirmed - Track: Posttraining & Midtraining - Room: Track 9 - Session type: session - Status: confirmed ## Session Description Verifiable rewards are the gold standard for RL training, but real-world agent tasks frequently lack clean deterministic evaluation objectives. This talk surveys our efforts to scale RL in non-verifiable settings -- including task synthesis, unsupervised environment design, and automatic judge calibration -- to ultimately enable self-improvement in production, grounded in real-world agent traces and domain-specific context. ## Media Evidence [Reinforcement Learning for Agents - Will Brown, ML Researcher at Morgan Stanley](https://www.youtube.com/watch?v=JIsgyk0Paic) (speaker-match related prior/adjacent AI Engineer video; captions: English auto-captions). - Source video: `youtube-JIsgyk0Paic` - Slide deck: [[youtube-JIsgyk0Paic-dense-slides|Dense Slides: Reinforcement Learning for Agents - Will Brown, ML Researcher at Morgan Stanley]] — 11 visible slide image(s); 11 HTML recreation(s).

### Speaker And Company Context
- [[will-brown|Will Brown]] — Researcher at [[prime-intellect|Prime Intellect]].

### Topics Covered
- [[agentic-search]]

### Derived Links And Source Material
- [[youtube-JIsgyk0Paic]] — related YouTube source page.
- [[youtube-JIsgyk0Paic-slides]] — slide evidence.
- [[youtube-JIsgyk0Paic-reconstructed-slides]] — slide evidence.
- [[youtube-JIsgyk0Paic-dense-slides]] — slide evidence.

### Novel Concepts / Clever Methods
- No highlighted novel concept has been detected yet.

### Evidence Boundary
This synthesis is based on the official schedule and linked source pages. It should be revisited when exact session recordings or transcript-backed secondary sources are available.
