SWE-bench Leaderboard: 13 Models and 4 Agents Tested Across Go, Java, Python, Rust, and TypeScript

13 Models and 4 Agents on SWE Tasks: Go, Java, Python, Rust, TS

SWE-bench Leaderboard: 13 Models and 4 Agents Tested Across Go, Java, Python, Rust, and TypeScript

I present a comprehensive leaderboard evaluating 13 AI models and 4 agents on software engineering tasks across Go, Java, Python, Rust, and TypeScript. The results highlight Fable 5 and Grok 4.5 as top performers in pass rates and cost efficiency. This analysis offers a clear view of current capabilities in automated code generation and debugging.

Fable 5 leads the pack with a 64.5% pass rate, proving that high-performance coding agents are now a reality.
  1. sathish316

    What does it mean when Fable 5 is 1st place and Opus 5 is 3rd place, while Claude code is 7th place? Which model and effort is used for Claude in 7th place, compared to 1st and 3rd?

  2. spullara

    They are all different problems for the different languages. I was hoping this was a benchmark that attempted to see which languages were more efficient to use with which models.

  3. dia80

    Why test Fable high effort vs Sol medium? Especially when Sol comes out 4-5x cheaper in their tests at those effort levels.

More from this day

2026-07-31