GEPA: Reflective Prompt Evolution Can Outperform
Reinforcement Learning
Chris Potts
https://www.youtube.com/watch?v=0bkwd9OYqfk
Model HotpotQA IFBench Hover PUPA Aggregate Improvement
Qwen3-8B
Baseline 42.33 36.90 35.33 80.82 48.85 —
MIPROv2 55.33 36.22 47.33 81.55 55.11 +6.26
GRPO 43.33 35.88 38.67 86.66 51.14 +2.29
GEPA 62.33 38.61 52.33 91.85 61.28 +12.44
My point here, though, is that both of them outperformed GRPO, which ought to be a kind of advanced RL-based post-training method, a fine-tuning method.
AI text/layout recreation from video frame; verify against source image.