The process ≠ answer story is not new
The Open Proof Corpus, Dekoninck and Petrov et. al. 2026
Correct Final Answer
Correct Proof
o3
Gemini-Pro
their reasoning often contains subtle logical errors masked by fluent language, posing significant risks for critical applications.
For every success story like the unit distance counterexample, there are likely thousands of pages generated for each of these problems, which have led nowhere.
an LLM agent with access to unit tests may delete failing tests rather than fix the underlying bug.
72% of reward hacking episodes include explicit chain-of-thought rationale, suggesting models often frame exploits as legitimate problem-solving.
AI text/layout recreation from video frame; verify against source image.