Cognition's SWE-2 Matches Fable 5.1 at 64% Lower Cost

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

Cognition's SWE-2 Matches Fable 5.1 at 64% Lower Cost

Cognition released SWE-2, a coding model that scores 50.0% on FrontierCode 1.1 Main—within one point of Fable 5.1 while being 64% cheaper. Post-trained from Kimi K3, it uses a novel RL algorithm that trains all reasoning-effort levels in one run by tuning cost penalties to the Pareto frontier's slope. SWE-2 also triples RL environments and stabilizes training with a length-weighted reward baseline. Available now in Devin Desktop and CLI.

Stronger engineering judgment allows the agent to write more complete solutions alongside fewer detours and redundant reads.
  1. postalcoder

    If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%).

    Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

  2. gruez

    Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked?

    https://www.youtube.com/watch?v=tNmgmwEtoWE

    As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in performance should be taken with a grain of salt.

  3. nullbio

    Where are the model stats? Is this open-weights? If not, why would I use this over DeepSeek Flash 4.1?

    I think these competing labs need to realize that no one wants another closed-weight model provider... We aren't even happy with the two we have right now, and their days are entirely numbered. If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper).

    The big labs love to release their new model and quantize after the first week. You don't have that problem using dirt cheap API rates on OpenRouter. DS 4.1 flash is also faster than fast mode Astra. OAI's subscription rates are good value, but now these new open-weight models are nearly as cheap on API usage rates. I honestly can't wait for the day we're not beholden to the two big labs anymore. No wonder there's so much fear pumping happening at the moment from Anthropic and their funded NGOs.

  4. captainregex

    I am skeptical. Lived experience is what matters and I don’t have anyone in my life (Devin shop) saying good things about SWE other than it’s free. Hope I’m wrong and it’s not so bad this time

  5. TheJCDenton

    > SWE-2 is post-trained from Kimi K3

    On the one hand I would have expected a completely new model, on the other hand it's an RL-ed K3 go Fable 5 capabilities, which demonstrate that this is probably possible, which is nice.

  6. pkilgore

    Not sure it matters when devin is the most consistently shit product I've used. And yes, I tried again, they wasted the money on the billboards.

  7. bobtheborg

    SWE 1.6 was great for small tasks. Very fast and good enough. 1.7 was unusable for me. Took more time thinking than GLM 5.2 and seemed to be generally running in circles. I tried it but abandoned it.

    Looking forward to 2 -- maybe it'll be usable

  8. mydreamof

    Seems like benchmaxing? For example for Terminal-Bench 4 it doesn't have great results. And why not show other benchmarks?

More from this day

2026-09-10