Cognition's SWE-2 tops Terminal-Bench 2.1 at 92.8, but trails frontier on long-horizon tasks
Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1

SWE-2 is Cognition's post-trained version of Kimi K3, a 2.8T-parameter MoE with 104B active per token. It scores 92.8 on Terminal-Bench 2.1 and 50.0 on FrontierCode 1.1 Main, just behind Fable 5.1 and GPT-6 Astra but at a claimed 64% lower cost. However, its 27.3 on Terminal-Bench 4.0 reveals a wide gap on long-horizon agentic work. Weights are proprietary, available only via Devin Desktop and CLI.
Terminal-Bench 2.1 is the highest number in the published table. The soft spot is Terminal-Bench 4.0, where SWE-2's 27.3 trails Fable 5.1 (55.8) and GPT-6 Astra (57.9) by a wide margin - long-horizon agentic work is where the gap to the frontier still lives.
- samusiam
Which is pretty much a useless (i.e., saturated, contaminated) benchmark now.
- forgot-my-pw
Terminal Bench 2/2.1 appears to be almost solved, so we probably shouldn't look too hard on that? Their Terminal Bench 4 score is soso.
Still quite impressive though.
- captainregex
my personal experience with swe has been…suboptimal. I am not sure how much I buy these benchmarks and it has a very “just blurt it out even if it’s probably not right” style but hey it’s free.