SWE-bench Leaderboard: 13 Models and 4 Agents Tested Across Go, Java, Python, Rust, and TypeScript
13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS

I present a comprehensive leaderboard evaluating 13 AI models and 4 agents on software engineering tasks across Go, Java, Python, Rust, and TypeScript. The results highlight Fable 5 and Grok 4.5 as top performers in pass rates and cost efficiency. This analysis offers a clear view of current capabilities in automated code generation and debugging.
Fable 5 leads the pack with a 64.5% pass rate, proving that high-performance coding agents are now a reality.
- sathish316
What does it mean when Fable 5 is 1st place and Opus 5 is 3rd place, while Claude code is 7th place? Which model and effort is used for Claude in 7th place, compared to 1st and 3rd?
- spullara
They are all different problems for the different languages. I was hoping this was a benchmark that attempted to see which languages were more efficient to use with which models.
- dia80
Why test Fable high effort vs Sol medium? Especially when Sol comes out 4-5x cheaper in their tests at those effort levels.