Markdown source

Dense Slides: WF2026: Software Factories & Keynotes ft. Microsoft, OpenAI, OpenClaw, Z.ai (GLM), MiniMax, HF

Source Video

WF2026: Software Factories & Keynotes ft. Microsoft, OpenAI, OpenClaw, Z.ai (GLM), MiniMax, HF

Method

This deck is slide-only. The existing captured video frame set supplies candidate frames, then local OpenCV rejects sponsor/title/speaker-only frames, crops visible slide surfaces, deduplicates, and saves the cropped slide images.

Cropped Visible Slides

slide-001.jpg

Slide text:

Nesled boogrs WNe

Loopcraft: The Art of Stacking Loops

Leel 6 / G

Tokens Turms Tools Tasks Automations

6 · software factories??????·ar gools, alocate. col tot nee. open eplorwtion

5.automatom·cron,soogonts,muRr-agem eitcoaorationad cepetisio:adarys

4 ·paal loop ·run,judga,rety eit:goal reathet 、huy

opent ealt vege off-goa. agent resu wogos me v

3·agmtturn·caf tooi,tedresun toit: no mare 10ol cafls: ≥ Nn5s

reas_(ite() →240803 rn_tests()

2·cat lpep·promgt. respond,aign eit Meofuf reney:oe hure

AH·BOAOH

user:“eplainrecursion* asistart: dreft reply humk甲 humpn. elgredrepsy

1. token loop· sample,appendl repiat tt: atop toetn·secondy

the cat tat

slide-002.jpg

Slide text:

# unedited, from the threads HARD DATA:

"different outputs at temperature = 0, "even at temp O you get different: r/LocaiLLaMA answers, you're using GPUs." mostly the MoE architecture." r/LocaiLLaMA 1,000 prompts temp 0. vLLM - Qwen-3-8B - 80 answers

"completely correct or completely wrong,. Hacker News depending on minute numerical diffs."

slide-003.jpg

Slide text:

# unedited, from the threads HARD DATA:

: r/locaiLlaMA "different outputs at temperature = 0, "even at temp O you get different answers, you're using GPUs.": r/LocaillaMA. mostly the MoE architecture." 1,000 prompts temp 0: vLLM: Qwen-3-8B - 80 answers

"completely correct or completely wrong.: Hacker News depending on' minute numerical diffs."

slide-004.jpg

Slide text:

It's throughput — how many cycles fit the same window.

7 cycles

143 cycles

slide-007.jpg

Slide text:

World's Fair AlEngincer

Are these fully-vibed pRs good?

Womantedroiunow

Were real companies, with real codebases, coding this way?

Were the end-to-end Al-generated PRs any good?

Gl Lab Vorld'sFair ALn In what ways did they fail?

Fair enAl groptio

ler. uFair Engineering the future of Al

Fair TADOG

slide-008.jpg

Slide text:

AI systems evolved faster than our evaluation methods

The Illusion The Reality

100% MAA Invisible Failure Hodes

75% Behavior Degraded Production

90% 50% 25% Reliability Gaps Unpredictable User

Benchmark Accuracy 0% T-0 T+10ms T+50ms T+10Gms

slide-009.jpg

Slide text:

Agenda

1. Meet OG Assist

2. The Origin Story

3. Betting on Effect

4. The Core Agent Loop

5. A2A, Evals & Sandboxing

6. Long Context Handling

7. Monitoring & Observability

8. Tools, Skills & Dev Workflows

slide-010.jpg

Slide text:

Origin Story

One bet on agents, one immediate yes.

slide-011.jpg

Slide text:

Example inport ( Chat. languagcHodel ) from "@elfect/si" Tho Hamess Our agent loop,

// Streaaing chat with tool suppart leport ( Etfect, Streaa ) fros "cffect" rebuilt

const streanirgchat = Effect.genifunction- () (

Const chat = yielo- chat.ctoty Effect-native.

Yielce chat

.stresnTexti(

proapt: "Generate a crcative story"

11

.pipeiStrean.runforEachiipartl => Effect.synci() = console.toglpart/l!)

The core loop started on LangGraph. It now runs fully

Effect-native typed control flow, structured concurrency, and

Source: https//effect-ts.github "zvt-o0- co- language model if we were to uh the agent loop. resource safety end to end, Aiso allows more granular control t

slide-012.jpg

Slide text:

The missing layer between evals and action

Observability

Your observability stack captures every tool call, every LLM completion, and every exception.

Evals

Your eval suite judges whether the final output was correct.

Agent

Context, Skills and .md file

The GAP

slide-013.jpg

Slide text:

In practice, most agent memory frameworks have focused on user continuity: preferences,

profile facts, conversation history, and long-lived personalization.

chat experiences is not self-improving learning system for production agents. 米

Approsch What It stores Retrleval signal Learns from outcomes?

ta?ut! Tulu iLryCher, hoet Ravr chat hlstory Recency No

Extracted facts, preteroncos Embedding similarity Ho

Tt tttg yiph L'ip'drieia. Entity relationshups over timo Graph travorsal + foconcy No

Verbal self-roflections Simlarity to current task Partialty: reflections capture ranked by outcome lossons, but rotrievat b not

agentRTx Ltlity scores Tast-lrkcd refoctions wth Simitasity weighted by outcome-derived utility Tm tehty tcorat updt sln it

slide-014.jpg

Slide text:

Memory as Reasoning

facts

User preferences

Reasoning

"Check settlement before Refund"

Static,

No Context

No history

Reranked based on usefulness

Context is updated based on task

learned from history

slide-015.jpg

Slide text:

AI Engineer World's Fair

Engineering the future of AI

slide-016.jpg

Slide text:

Room to act, safely.

Code runs in isolated sandboxes, so agents can take real action without putting production systems or customer data at risk.

slide-017.jpg

Slide text:

The Paradigm Shift: Output vs. Behavior

Traditional LLM Evaluation Agent Evaluation

Goal Output Accuracy Workflow Behavior

Environment Static Datasets Dynamic Contexts

Execution Single-path Processing 一 Multi-path & Tool Dependent

Failure Mode Hallucination Cascading Workflow Failure

slide-018.jpg

Slide text:

Agents keep failing at the same tasks.

Gartner's 2025 AI deployment survey found that 85% of AI projects fail in production. McKinsey's 2025 State of AI report found that fewer than 20% of AI pilots scale to production within 18 months.

slide-019.jpg

Slide text:

World's Fair AlEngineer BEWARE' AI OF POLICE deflim (q: Question) (v: IsProper q): IO (∑ (a: Answer), IsSafeText q v a

World'sFair BEWARE' AI OF. PLAN def Ilm (q: Question) (v: IsProper q): (∑ (a: IO Answer),IsSafelO q v a

Browst.

World's In Code They Act, In Proof We Trust

Erik Meijer / Research Scholar Leibniz Labs

Hidden Non-Slide Evidence

Classification audit: raw/sources/slide-ai-classification/dense/htM02KMNZnk/audit.json