Markdown source

Reinforcement Learning without Verifiable Rewards

Conference Context

Session Description

Verifiable rewards are the gold standard for RL training, but real-world agent tasks frequently lack clean deterministic evaluation objectives. This talk surveys our efforts to scale RL in non-verifiable settings -- including task synthesis, unsupervised environment design, and automatic judge calibration -- to ultimately enable self-improvement in production, grounded in real-world agent traces and domain-specific context.

Media Evidence

Reinforcement Learning for Agents - Will Brown, ML Researcher at Morgan Stanley (speaker-match related prior/adjacent AI Engineer video; captions: English auto-captions).

Evidence Graph

This evidence graph is generated from currently linked source material: official schedule text, related video pages, cached transcripts, visible slide text, dense/reconstructed slide pages, and AI slide-classification audits.

Media Signals

Agent Reading Notes

Use these signals to refine the synopsis, topic links, people/company context, and method notes. If a source is a related external video rather than an exact official recording, keep it framed as supporting evidence.

Transcript Status

Related video transcript availability: English auto-captions. Treat this as supporting context, not a recording of this exact scheduled session unless later confirmed. Not fetched yet.

People

Supporting Slides

Slide Evidence

Synthesis

Synthesized Breakdown

Reinforcement Learning without Verifiable Rewards ## Conference Context - Date/time: 2026-06-30 · 1:30pm-1:50pm - Track/room: Posttraining & Midtraining · Track 9 - Speaker(s): Will Brown - Session type/status: session · confirmed - Track: Posttraining & Midtraining - Room: Track 9 - Session type: session - Status: confirmed ## Session Description Verifiable rewards are the gold standard for RL training, but real-world agent tasks frequently lack clean deterministic evaluation objectives. This talk surveys our efforts to scale RL in non-verifiable settings -- including task synthesis, unsupervised environment design, and automatic judge calibration -- to ultimately enable self-improvement in production, grounded in real-world agent traces and domain-specific context. ## Media Evidence Reinforcement Learning for Agents - Will Brown, ML Researcher at Morgan Stanley (speaker-match related prior/adjacent AI Engineer video; captions: English auto-captions). - Source video: youtube-JIsgyk0Paic - Slide deck: Dense Slides: Reinforcement Learning for Agents - Will Brown, ML Researcher at Morgan Stanley — 11 visible slide image(s); 11 HTML recreation(s).

Speaker And Company Context

Topics Covered

Derived Links And Source Material

Novel Concepts / Clever Methods

Evidence Boundary

This synthesis is based on the official schedule and linked source pages. It should be revisited when exact session recordings or transcript-backed secondary sources are available.