Earlier quoted context omitted.
This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.
Benchmaxxing is the default case, and always has been. It's really, really difficult to avoid it even when you care to stop yourself; and it's not even just a problem in machine learning, it's the standard failure mode of all minds capable of learning, human, animal, artificial. Even pure genetics has this problem. Viruses and cancers also demonstrate this behaviour, with the bench being evolution's only option: repr…
Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
121–130 of 201 posts
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#122Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#123Earlier quoted context omitted.
> If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper). Only in the world where the incumbents don't react. Eg if they saw lots of users moving away, they'd drop prices or do something else.
1M cached tokens on deepseek is $0.006, the big labs can't sell anywhere close to this, they have funders expecting returns and huge overhead. btw I've had a ton of fun with the new deepseek today, I was waiting for my OpenAI 5h limit reset and decided to give it some problems for fun, got pretty great results. Tried some harder problems and still got great results. I don't expect it to be Sol class or anything but I…
I'm willing to tolerate babysitting things a lot more if I know I'll get almost instant results.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#124Earlier quoted context omitted.
> If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper). Only in the world where the incumbents don't react. Eg if they saw lots of users moving away, they'd drop prices or do something else.
1M cached tokens on deepseek is $0.006, the big labs can't sell anywhere close to this, they have funders expecting returns and huge overhead. btw I've had a ton of fun with the new deepseek today, I was waiting for my OpenAI 5h limit reset and decided to give it some problems for fun, got pretty great results. Tried some harder problems and still got great results. I don't expect it to be Sol class or anything but I…
But great that we have a new leader in performance/price in that segment.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#125Earlier quoted context omitted.
> If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper). Only in the world where the incumbents don't react. Eg if they saw lots of users moving away, they'd drop prices or do something else.
OpenAI just paused new subscriptions to their $200 plan. They are in a rock and a hard place. Obviously the Astras and Fables of the world are exponentially more expensive, but for...less than exponential returns. The question is whether they can leverage the marginal advantage into something that justifies the diminishing returns before the bottom catches up to them. On the one hand you, if you bought a lot of compu…
That's like four years of ChatGPT + Claude subscription.
Eight years if only ChatGPT, or sixteen years of the Pro 5x subscription.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#126If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
Those 27.3% are still in the ballpark of modern models: - Sonnet 5 - 12.4% - Luna - 17.3% - Grok 4.6 - 20.3% - Sol - 37.3% - GLM 5.3 - 41.8% - Opus 5 - 51.8%
Also a lot of questions to benchmark because opus 5 is completely useless model right now.
I think that the main problem with opus that they try to solve context size optimization problem, and that is the main reason why it speaks like alien with only one technical dictionary at hand. So why it is so good?
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#127Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked? https://www.youtube.com/watch?v=tNmgmwEtoWE As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in perfor…
A friend recently pushed me to try out their coding platform, Devin, after I decided to move away from Cursor. I had the same reaction: "What, the con artists from like 2024?" But after some cajoling, I gave it a shot and was pleasantly surprised. I guess they learned their lessons, grew up, and are doing good work now, maybe?
They seem to love a good overpromise.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#128Why doesn't clickbait trash like this get moderated ?
Dang is pretty good at enforcing stuff. There's a flag button and you can email reports if you're really bothered.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#129Earlier quoted context omitted.
please keep thinking this so i can relax with my automated job
It's not you, it's X... but what would you expect of a nepo-baby economy of little swines. This is like the nepo wet-dream on steroids. Incompetence and delulu
Watching the AI slop my sales reps put in their emails is disgusting but the reply telling them how great of a job they are doing and how insightful their email was says differently.
Many people are laughing to the bank while you are still running `--help` to figure out how to run a complex command.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#130Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked? https://www.youtube.com/watch?v=tNmgmwEtoWE As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in perfor…
A friend recently pushed me to try out their coding platform, Devin, after I decided to move away from Cursor. I had the same reaction: "What, the con artists from like 2024?" But after some cajoling, I gave it a shot and was pleasantly surprised. I guess they learned their lessons, grew up, and are doing good work now, maybe?