Live data from Hacker News

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

cognition.com

121–130 of 201 posts

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#121
post #115

Earlier quoted context omitted.

This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.

Benchmaxxing is the default case, and always has been. It's really, really difficult to avoid it even when you care to stop yourself; and it's not even just a problem in machine learning, it's the standard failure mode of all minds capable of learning, human, animal, artificial. Even pure genetics has this problem. Viruses and cancers also demonstrate this behaviour, with the bench being evolution's only option: repr…

Why would I adapt myself to unseen environments unnecessarily?

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#122

Earlier quoted context omitted.

I have not met a single person/company that uses Devin… does anyone here actually use it?

my friends at infosys are being trained on it

Deeply unserious company so not surprising

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#123
post #27

Earlier quoted context omitted.

> If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper). Only in the world where the incumbents don't react. Eg if they saw lots of users moving away, they'd drop prices or do something else.

1M cached tokens on deepseek is $0.006, the big labs can't sell anywhere close to this, they have funders expecting returns and huge overhead. btw I've had a ton of fun with the new deepseek today, I was waiting for my OpenAI 5h limit reset and decided to give it some problems for fun, got pretty great results. Tried some harder problems and still got great results. I don't expect it to be Sol class or anything but I…

Yep. 4.1 Flash is good enough for most routine coding things, but it also makes up for a lot of weakness by being so fast (and cheap of course).

I'm willing to tolerate babysitting things a lot more if I know I'll get almost instant results.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#124
post #27

Earlier quoted context omitted.

> If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper). Only in the world where the incumbents don't react. Eg if they saw lots of users moving away, they'd drop prices or do something else.

1M cached tokens on deepseek is $0.006, the big labs can't sell anywhere close to this, they have funders expecting returns and huge overhead. btw I've had a ton of fun with the new deepseek today, I was waiting for my OpenAI 5h limit reset and decided to give it some problems for fun, got pretty great results. Tried some harder problems and still got great results. I don't expect it to be Sol class or anything but I…

That's 1/3 of GPT 5.6 Luna. It seems rather close to me.

But great that we have a new leader in performance/price in that segment.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#125
post #27

Earlier quoted context omitted.

> If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper). Only in the world where the incumbents don't react. Eg if they saw lots of users moving away, they'd drop prices or do something else.

OpenAI just paused new subscriptions to their $200 plan. They are in a rock and a hard place. Obviously the Astras and Fables of the world are exponentially more expensive, but for...less than exponential returns. The question is whether they can leverage the marginal advantage into something that justifies the diminishing returns before the bottom catches up to them. On the one hand you, if you bought a lot of compu…

I could buy four of them for ~20k.

That's like four years of ChatGPT + Claude subscription.

Eight years if only ChatGPT, or sixteen years of the Pro 5x subscription.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#126

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

Those 27.3% are still in the ballpark of modern models: - Sonnet 5 - 12.4% - Luna - 17.3% - Grok 4.6 - 20.3% - Sol - 37.3% - GLM 5.3 - 41.8% - Opus 5 - 51.8%

GLM 5.3 looks strange, because of this Chinese labs benchmaxx moto. So rather they have emergent abilities or...

Also a lot of questions to benchmark because opus 5 is completely useless model right now.

I think that the main problem with opus that they try to solve context size optimization problem, and that is the main reason why it speaks like alien with only one technical dictionary at hand. So why it is so good?

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#127
post #35

Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked? https://www.youtube.com/watch?v=tNmgmwEtoWE As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in perfor…

A friend recently pushed me to try out their coding platform, Devin, after I decided to move away from Cursor. I had the same reaction: "What, the con artists from like 2024?" But after some cajoling, I gave it a shot and was pleasantly surprised. I guess they learned their lessons, grew up, and are doing good work now, maybe?

When they launched Devin it was supposedly at the performance of an engineering intern. Friends who used it found the bad parts of an intern (tons of handholding, review required) but it didn’t learn from mistakes or add throughput.

They seem to love a good overpromise.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#128

Why doesn't clickbait trash like this get moderated ?

Because it's not clickbait trash / doesn't violate the site rules?

Dang is pretty good at enforcing stuff. There's a flag button and you can email reports if you're really bothered.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#129

Earlier quoted context omitted.

please keep thinking this so i can relax with my automated job

It's not you, it's X... but what would you expect of a nepo-baby economy of little swines. This is like the nepo wet-dream on steroids. Incompetence and delulu

If you are too dumb or too lazy to figure out what this guy has, then you are the problem and will be looked at as a relic.

Watching the AI slop my sales reps put in their emails is disgusting but the reply telling them how great of a job they are doing and how insightful their email was says differently.

Many people are laughing to the bank while you are still running `--help` to figure out how to run a complex command.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#130
post #35

Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked? https://www.youtube.com/watch?v=tNmgmwEtoWE As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in perfor…

A friend recently pushed me to try out their coding platform, Devin, after I decided to move away from Cursor. I had the same reaction: "What, the con artists from like 2024?" But after some cajoling, I gave it a shot and was pleasantly surprised. I guess they learned their lessons, grew up, and are doing good work now, maybe?

Not all that surprising. The original Devin really was just an early attempt at agentic coding before models were really even trained for it. Now that it's a well established pattern and we've figured out what works, I'm not surprised they've morphed into something reasonble.
Post reply on HN