Slides: SWE-rebench: Lessons from Evaluating Coding Agents — Ibragim Badertdinov, Nebius
Source Video
SWE-rebench: Lessons from Evaluating Coding Agents — Ibragim Badertdinov, Nebius
Relationship To World's Fair 2026
These slides are extracted from a public AI Engineer YouTube video connected to World's Fair 2026. Speaker-matched clips are supporting context unless later confirmed as exact session recordings; official livestream recordings are day-level/event-level source material.
Related Scheduled Sessions
- No individual scheduled session mapping has been assigned yet; treat this as an event livestream deck.
Extracted Slides

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
title_cardconfidence0.91 - Text source: agent_vision.
Slide text:
SWE-rebench: Lessons from Evaluating Coding Agents on Real Software Engineering Tasks

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.99 - Text source: agent_vision.
Slide text:
Why evals matter now?
Models improved. Choosing became harder.
Vibe checks do not scale
SWE performance grows rapidly
Options change every month
* SWE-rebench 2026_02 tasks

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.96 - Text source: advanced OCR
rapidocr-live/bright-screen/opencv-adaptive. - OCR decision: ready — multi-column slide with screenshots and small embedded labels
Slide text:
Anatomy of a Task: A Task Is More Than Text
Task description - original issue text Regresslon causod by changes for woakref ot fllesystom 1284
★ AIE +
Sandbox environment -
executable Docker image shs256141e015c4.:9
Verifier - tests from the PR. FAIL_TO PASS + PASS TO PASS Sire 14GB Leut updaitd 19 d3y2 a90
docker dull s=ereb+nch/seb.eva]
-dev.1776.pyfakets-1286
AlEngineer
EUROPE

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.97 - Text source: advanced OCR
rapidocr-live/bright-screen/opencv-adaptive. - OCR decision: ready — code screenshot and small text in a structured slide layout
Slide text:
What Makes a Good Task?
Problem description
1. A good task balances clarily
+ ★ AIE Reliable verifier 3. 2. Too easy and too hard both fail Complexity: breadth or depth 35 Gettest listaansles_oapty_or eone_obiectsicbjccts_vatue] Colcl1ent coch,Moc Spl cllent.It_cuthent1eated.retrn Vluo o True Col ellent.coject tore.tnt,retum vluo o? ebjectb objects_wlue. Gxcebrto
1. Should reward actual fixes, should reject fake
solutions (nltlaixe_coa:simertapl ctient_to_ustaol_clfent]
2. Not too narrow, not too wide rewtt oClRrnerh1mvkellem, Gcloud,Cobjet-store
Stable infrastructure is part of eval → Sseerel ho cojects tound atl'test 'in resutt.cot pat
1. Minimal infra noise during runs + contaier.aplclint.object_store.list.asert"called_onc_u
2. ( Connection might blink, images might become
stale. pipeline might break (1970s bug)
AIEngineer
EUROPE

- Recreated text/layout view: open HTML recreation
- AI slide classifier:
content_slideconfidence0.96 - Text source: advanced OCR
rapidocr-live/bright-screen/contrast. - OCR decision: ready — small command list and dense layout on the right side
Slide text:
Execution Setup: Minimal Agent, Strong Infrastructure
OPEN
grep
Minimalistic agent (open, edit, bash)
AIE YOLO setup ReAct + demo - tools + no_demo Agent<Infrastructure python python3 EDIT find git REPLACE GOTO cat cd
· Claude-Opus-4.6 top commands from our agent sed SUBHIT 1s SCROLL_DOAN rm
# Braintrust WorkOs OpenAi
Classification audit: raw/sources/slide-ai-classification/slides/wcUJWP6WpGM/audit.json
Slide-Derived Subjects To Review
Subject extraction uses video title, related session titles/descriptions, transcript context, and OCR text when available. OCR is best-effort and should be reviewed against the embedded slide images.