Free stealth model big-pickle tops SWE Atlas Codebase QnA leaderboard among open scaffolds

Big Pickle on SWE Atlas – Codebase QnA

OpenCode Zen's free stealth model, big-pickle, scored 50.8% on Scale AI's SWE Atlas Codebase QnA benchmark, outperforming all official leaderboard entries using the same mini-swe-agent scaffold, including GLM 5.2 and GPT-5.6-Sol. Only Claude models on their native scaffold scored higher. The run used a single trial and reduced sandbox resources, with caveats about model identity and data exposure. Full results and reproduction configs are provided.

Within the Mini-SWE-Agent scaffold class — the apples-to-apples comparison — this run outscores every entry on the official leaderboard, and it also tops the Codex-scaffold GPT entries.

More from this day

2026-08-16