Benchmarking Cline w/ updated prompt on SWE-Bench Lite
GPT-4.1
Claude Sonnet 4.5
~15%
Improvement, just through rules
No fine-tuning, no tool changes, no architecture changes. JUST RULES.
GPT-4.1 achieved performance near Sonnet 4.5, which is widely considered state of the art for coding questions
2/3 cost!
GPT-4.1 achieved performance near Sonnet 4.5, which is widely considered state of the art for coding questions
2/3 cost!