Live data from Hacker News

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

cognition.com

41–50 of 201 posts

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#41

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

Those 27.3% are still in the ballpark of modern models:

- Sonnet 5 - 12.4%

- Luna - 17.3%

- Grok 4.6 - 20.3%

- Sol - 37.3%

- GLM 5.3 - 41.8%

- Opus 5 - 51.8%

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#42
post #27
post #12

Where are the model stats? Is this open-weights? If not, why would I use this over DeepSeek Flash 4.1? I think these competing labs need to realize that no one wants another closed-weight model provider... We aren't even happy with the two we have right now, and their days are entirely numbered. If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the cl…

> If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper). Only in the world where the incumbents don't react. Eg if they saw lots of users moving away, they'd drop prices or do something else.

1M cached tokens on deepseek is $0.006, the big labs can't sell anywhere close to this, they have funders expecting returns and huge overhead.

btw I've had a ton of fun with the new deepseek today, I was waiting for my OpenAI 5h limit reset and decided to give it some problems for fun, got pretty great results. Tried some harder problems and still got great results. I don't expect it to be Sol class or anything but I really didn't expect it to be anywhere near this good so we'll see where it ends up. And it's really fun throwing crazy amount of tokens at the wall for ~free instead of watching the subscription limits tick closer while your agents churn away.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#43
post #12

Where are the model stats? Is this open-weights? If not, why would I use this over DeepSeek Flash 4.1? I think these competing labs need to realize that no one wants another closed-weight model provider... We aren't even happy with the two we have right now, and their days are entirely numbered. If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the cl…

This really just exists so cognition can stop spending API tokens with Anthropic or OpenAI.

Basically any successful AI based service will do this because at scale the frontier models are expensive and you’ll have enough data to fine tune your own.

Same reason Harvey is doing models now and basically every other provider

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#44
post #23

Earlier quoted context omitted.

Yeah, this echoes my thoughts. I will be very surprised if a model with 2.8T parameters reaches the intelligence and capabilities of 10T parameter models. RL can take things far, but not that far.

Closed weights AND benchmaxxed. Somehow this company raised 2bil at a 48bil valuation. Pure insanity. I feel bad for their investors (not really, but... Still). Andreessen Horowitz is being played like a fiddle.

> Andreessen Horowitz is being played like a fiddle.

Andreessen Horowitz is not being played like a fiddle here. This might be their only investment in a decade that isn’t entirely predicated on being a scam.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#45
post #12

Where are the model stats? Is this open-weights? If not, why would I use this over DeepSeek Flash 4.1? I think these competing labs need to realize that no one wants another closed-weight model provider... We aren't even happy with the two we have right now, and their days are entirely numbered. If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the cl…

> we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper).

DS 4 Flash requires large amounts of memory to run at reasonable quants (I think a system with 160 GB or so). DS 4.1 Flash is even larger, I think around 250 GB.

Any DS version is dumb when compared (in realworld tasks) to Astra/Opus 5, which means, one would spend thousands of dollars, and still need to rely on cloud services to do jobs that are non trivial.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#46

At work I setup a cloud worker, where i can spin up as many concurrent agents I want, with unlimited fable 5.1 (thanks employer!!). I now just work from my phone, and speak into the agents as they run. I dont write code and I dont write documents. I work on very complicated distributed systems. I dont open my laptop most days. Its a legacy brick I carry around. Some of my coworkers are still doing things by hand, and…

that's why after 1 year of product development of these AI 20x maxxed speed, we reached AGI 'wizards', there's really no difference in output, outstanding bugs no longer get solved and sites still suck, even doing things that were just regular development 20 years ago. Are you sure they aren't only producing 2.5% of your output that you manage just by farting into your phone? Are you sure it's 25% really? Seems way t…

please keep thinking this so i can relax with my automated job

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#47
post #40

Earlier quoted context omitted.

[flagged]

A lot of people seem to be the Antichrist these days. I’d have though we’d get a lull after the millennium but it’s all the rage.

I guess we also need more antipopes.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#48
post #40

Earlier quoted context omitted.

[flagged]

A lot of people seem to be the Antichrist these days. I’d have though we’d get a lull after the millennium but it’s all the rage.

Well, the antichrist should be in and around this AI thing for one particular reason:

The devil cannot create anything of his own because he is not God, by definition. We have already observationally defined generative AI as something that cannot create anything novel in the sense it cannot output anything it has never seen (cannot create new, always a re-assortment of what is).

In that way , AI is a perfect mimicry of how the devil operates (in totality, as the devil perverts and replicates anything good, often subtly and always deceptively), which is to thieve off God, steal.

So he would be around, if you catch my drift, right about now. And I wouldn’t be shocked if he’s on HN, and that he would chose technology as the vessel. And ultimately, when it’s all said and done, I would not be shocked that those who studied and developed AI, did so for the devil whether they were aware or not.

Anyway, let a poor Christian have his end-times hypothesis.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#49
post #14
post #13

Earlier quoted context omitted.

With the Devin subscription even at the 20$ plan, they offered unlimited SWE 1.7 usage. Wondering if they do the same for SWE 2.

SWE-2 is free for all subscribers on the CLI to try out for the next month :)

What about the gui/windsurf app? Same as cli?

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#50

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.

> Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"?

A model that was released a couple months ago scores 50% higher than SWE-2, a model released today, on an out-of-sample benchmark. Can I say I’ve come out of this more impressed with Sol?

Like you said, TB2 is saturated. Nobody would bat an eyelash at 90%. And yet here comes SWE-2 coming off top rope with an emphatic 92.4%. this is the definition of bench maxxing.

Post reply on HN