The Paradigm Shift: Output vs. Behavior
Traditional LLM Evaluation
Agent Evaluation
Goal
Environment
Execution
Failure Mode
Output Accuracy
Static Datasets
Single-path Processing
Hallucination
Workflow Behavior
Dynamic Contexts
Multi-path & Tool Dependent
Cascading Workflow Failure
AI text/layout recreation from video frame; verify against source image.