Live data from Hacker News

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

cognition.com

1–10 of 201 posts

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#6

Seems like benchmaxing? For example for Terminal-Bench 4 it doesn't have great results. And why not show other benchmarks?

Probably, FrontierCode is made by Cognition itself. The model also seems worse in every way than DeepSeek v4.1 Flash, launched today.

Also the submitter's account is very new which makes me suspicious of self-promotion.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#7
As an Econ graduate, pretty cool seeing Pareto in the "AI-bro" zeitgeist. Slightly surreal watching a 1906 welfare economics idea get rediscovered as a plotting convention. The original, if anyone fancies 579 pages of Italian: https://archive.org/details/manualedieconomi00pareuoft. There is an English translation somewhere.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#8
If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%).

Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#9
post #4

"SWE-2 is post-trained from Kimi K3"

Yeah, I'd expect model performance to be super spiky on SWE work, at least they admit it with the name of the model. It's distilled from an already-distilled model.

Maybe still worth it if their "64% cheaper" figure holds.

Post reply on HN