---
title: "Loophole - Adversarial Agents To Stress Test Your Morality"
category: "talks"
date: "2026-07-01"
time: "1:30pm-1:50pm"
track: "Harness Engineering"
room: "Main Stage"
speakers: ["Brendan Rappazzo"]
sourceLabels: ["Official conference schedule", "Public YouTube metadata"]
scheduleTrack: "Harness Engineering"
scheduleRoom: "Main Stage"
scheduleLabels: ["Harness Engineering", "Main Stage", "session", "confirmed"]
---
# Loophole - Adversarial Agents To Stress Test Your Morality

## Conference Context
- Date/time: 2026-07-01 · 1:30pm-1:50pm
- Track/room: Harness Engineering · Main Stage
- Speaker(s): Brendan Rappazzo
- Session type/status: session · confirmed

- Track: Harness Engineering
- Room: Main Stage
- Session type: session
- Status: confirmed

## Session Description
Most natural language specifications have holes their authors didn't notice - and writing more rules tends to create more holes. I built Loophole to try a different approach: point adversarial agents at a spec until it stops breaking. You give the system a set of natural language principles. An AI drafts a formal codified version. Two adversarial agents go to work - one finds cases the code permits but the principles forbid, the other finds cases the code forbids but the principles allow. A judge agent patches the code when it can, but only if the fix doesn't contradict any prior ruling. When a contradiction can't be resolved, it escalates to you. Every decision becomes binding precedent, so the constraint space tightens round after round. I started with moral and legal reasoning as the demo, and on its own that's already interesting - it turns into a kind of game where you discover contradictions in your own beliefs that you didn't know were there. But the pattern generalizes well past that. The same loop works for company policies that need to survive contact with edge cases. For making chatbot system prompts adversarially robust. For stress-testing eval rubrics. And, taking the long view, for something like a smarter legislative process - where proposed laws get checked against the public's stated values before they pass, and the contradictions surface before they hit a courtroom. The talk walks through how the harness works, the design choices that matter (especially why precedent is the load-bearing piece), what kinds of specs it handles well, where it breaks, and what it would take to push it further. All code is open source.

## Media Evidence
No related AI Engineer channel video found yet.

## Evidence Graph
This evidence graph is generated from currently linked source material: official schedule text, related video pages, cached transcripts, visible slide text, dense/reconstructed slide pages, and AI slide-classification audits.

### Media Signals
No linked video, transcript, or slide source has been attached yet.

### Agent Reading Notes
Use these signals to refine the synopsis, topic links, people/company context, and method notes. If a source is a related external video rather than an exact official recording, keep it framed as supporting evidence.

## Transcript Status
No official session recording transcript was found by exact title match on the AI Engineer YouTube channel during this run.

## People
- [[brendan-rappazzo]]

## Notes
- Pending transcript synthesis when an official recording or confirmed matching video is available.

## Synthesis
### Synthesized Breakdown
# Loophole - Adversarial Agents To Stress Test Your Morality ## Conference Context - Date/time: 2026-07-01 · 1:30pm-1:50pm - Track/room: Harness Engineering · Main Stage - Speaker(s): Brendan Rappazzo - Session type/status: session · confirmed - Track: Harness Engineering - Room: Main Stage - Session type: session - Status: confirmed ## Session Description Most natural language specifications have holes their authors didn't notice - and writing more rules tends to create more holes. I built Loophole to try a different approach: point adversarial agents at a spec until it stops breaking. You give the system a set of natural language principles. An AI drafts a formal codified version.

### Speaker And Company Context
- [[brendan-rappazzo|Brendan Rappazzo]] — Machine Learning Scientist at [[morgan-stanley|Morgan Stanley]].

### Topics Covered
- [[agent-security]]
- [[agentic-search]]
- [[coding-agents]]

### Derived Links And Source Material

### Novel Concepts / Clever Methods
- No highlighted novel concept has been detected yet.

### Evidence Boundary
This synthesis is based on the official schedule and linked source pages. It should be revisited when exact session recordings or transcript-backed secondary sources are available.
