Live data from Hacker News

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

cognition.com

21–30 of 201 posts

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#21

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"?

Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#22
post #11

Please correct me if I'm wrong, but this appears to require Devin to use? I'm disappointed to see I need to use a bespoke platform to interact with this agent, to the point that I probably won't be trying it.

SWE-2 is free to use for users like yourself for the next month, and almost all usage should be supported via our CLI ( https://docs.devin.ai/cli ) :) Disclaimer: I work at Cognition, although was not involved in SWE-2

But I don't want to use your CLI. I already have my own harnesses and workflows. The friction is too high to "just try out" a new model like this. It would be preferable if I can evaluate it over, say, open router like all the other models and then decide from there if it's worth downloading a bespoke tool chain for only 1 lab's models

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#23

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

Yeah, this echoes my thoughts. I will be very surprised if a model with 2.8T parameters reaches the intelligence and capabilities of 10T parameter models. RL can take things far, but not that far.

Closed weights AND benchmaxxed. Somehow this company raised 2bil at a 48bil valuation. Pure insanity. I feel bad for their investors (not really, but... Still). Andreessen Horowitz is being played like a fiddle.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#24

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

Yeah, this echoes my thoughts. I will be very surprised if a model with 2.8T parameters reaches the intelligence and capabilities of 10T parameter models. RL can take things far, but not that far.

[deleted]

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#25

At work I setup a cloud worker, where i can spin up as many concurrent agents I want, with unlimited fable 5.1 (thanks employer!!). I now just work from my phone, and speak into the agents as they run. I dont write code and I dont write documents. I work on very complicated distributed systems. I dont open my laptop most days. Its a legacy brick I carry around. Some of my coworkers are still doing things by hand, and…

that's why after 1 year of product development of these AI 20x maxxed speed, we reached AGI 'wizards', there's really no difference in output, outstanding bugs no longer get solved and sites still suck, even doing things that were just regular development 20 years ago. Are you sure they aren't only producing 2.5% of your output that you manage just by farting into your phone? Are you sure it's 25% really? Seems way to high, days when I have diarrhoea my AI agents move even faster

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#26
I like Cognition as a company and hope they succeed. Seemingly excellent engineering org.

I used to really like Windsurf. (Now Devin. Kind of? But also now Antigravity.) I still use it as my editor but haven't touched the agent for a while simply due to the rise of Codex.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#27
post #12

Where are the model stats? Is this open-weights? If not, why would I use this over DeepSeek Flash 4.1? I think these competing labs need to realize that no one wants another closed-weight model provider... We aren't even happy with the two we have right now, and their days are entirely numbered. If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the cl…

> If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper).

Only in the world where the incumbents don't react. Eg if they saw lots of users moving away, they'd drop prices or do something else.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#28
post #4

"SWE-2 is post-trained from Kimi K3"

Yeah, I'd expect model performance to be super spiky on SWE work, at least they admit it with the name of the model. It's distilled from an already-distilled model. Maybe still worth it if their "64% cheaper" figure holds.

I don't think you know what distill means

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#29
post #23

Earlier quoted context omitted.

Yeah, this echoes my thoughts. I will be very surprised if a model with 2.8T parameters reaches the intelligence and capabilities of 10T parameter models. RL can take things far, but not that far.

Closed weights AND benchmaxxed. Somehow this company raised 2bil at a 48bil valuation. Pure insanity. I feel bad for their investors (not really, but... Still). Andreessen Horowitz is being played like a fiddle.

The Cursor acquisition shows that it’s possible for these valuations to be justified. But Cursor was more successful and bent the truth much less.

While I wouldn’t expect anything good for Cognition’s fate, it’s a much safer bet than Thinking Machines, SSI, and some others.

Though they’ll be in big trouble if the more talented Chinese labs stop letting them repackage their work.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#30

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.

Is Sol benchmaxxed? Of course it is. Altman was caught in previous attempts trying to game benchmarks, does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with?
Post reply on HN