Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
1–10 of 201 posts
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#2Seems like benchmaxing? For example for Terminal-Bench 4 it doesn't have great results. And why not show other benchmarks?
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#3SWE-1.5 was surprisingly good when I used it last. I feel like Cognition is one of the solid players that’s flying a bit under the radar while Anthropic and OpenAI race to IPO.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#4"SWE-2 is post-trained from Kimi K3"
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#5Please correct me if I'm wrong, but this appears to require Devin to use? I'm disappointed to see I need to use a bespoke platform to interact with this agent, to the point that I probably won't be trying it.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#6Seems like benchmaxing? For example for Terminal-Bench 4 it doesn't have great results. And why not show other benchmarks?
Probably, FrontierCode is made by Cognition itself. The model also seems worse in every way than DeepSeek v4.1 Flash, launched today.
Also the submitter's account is very new which makes me suspicious of self-promotion.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#7As an Econ graduate, pretty cool seeing Pareto in the "AI-bro" zeitgeist. Slightly surreal watching a 1906 welfare economics idea get rediscovered as a plotting convention. The original, if anyone fancies 579 pages of Italian: https://archive.org/details/manualedieconomi00pareuoft. There is an English translation somewhere.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#8If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%).
Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#9"SWE-2 is post-trained from Kimi K3"
Yeah, I'd expect model performance to be super spiky on SWE work, at least they admit it with the name of the model. It's distilled from an already-distilled model.
Maybe still worth it if their "64% cheaper" figure holds.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#10> SWE-2 is post-trained from Kimi K3
On the one hand I would have expected a completely new model, on the other hand it's an RL-ed K3 go Fable 5 capabilities, which demonstrate that this is probably possible, which is nice.