Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
101–110 of 201 posts
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#102Where are the model stats? Is this open-weights? If not, why would I use this over DeepSeek Flash 4.1? I think these competing labs need to realize that no one wants another closed-weight model provider... We aren't even happy with the two we have right now, and their days are entirely numbered. If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the cl…
> If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper). Only in the world where the incumbents don't react. Eg if they saw lots of users moving away, they'd drop prices or do something else.
On the one hand you, if you bought a lot of compute a couple years ago (perceived demand, perceived shortage) you are in a good spot temporarily. But the counter to that is that everyone else is becoming more compute efficient so maybe that advantage isn't what people thought it would be. I can almost, almost run DS4.1 Flash at home. 4 sparks can do it at 200+ tokens per second. I have two Sparks, so I am not in the club. Neither is your average laptop owner or gamer either. But your average HN software engineer can probably easily swing 2 sparks.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#103Not sure it matters when devin is the most consistently shit product I've used. And yes, I tried again, they wasted the money on the billboards.
I have not met a single person/company that uses Devin… does anyone here actually use it?
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#104Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#105Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked? https://www.youtube.com/watch?v=tNmgmwEtoWE As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in perfor…
A friend recently pushed me to try out their coding platform, Devin, after I decided to move away from Cursor. I had the same reaction: "What, the con artists from like 2024?" But after some cajoling, I gave it a shot and was pleasantly surprised. I guess they learned their lessons, grew up, and are doing good work now, maybe?
sorry so many buzzwords to say, the capabilities to do this kind of work are more accessible and easier to manage, so now it works!
Good to see, and agree they were severely overhyping their product back then.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#106https://cognition.com/frontiercode
Which is too bad, since all of the gains here appear to be from massively reduced output tokens?
The model SWE-2 is based on, Kimi K3, is cheaper per token than Sol, but costs more per task (ArtificialAnalysis) due to using way more tokens.
Whereas, based on the graphs, SWE-2 appears even more token-efficient than Sol! That might have been worth showing off, if true.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#107Earlier quoted context omitted.
A lot of people seem to be the Antichrist these days. I’d have though we’d get a lull after the millennium but it’s all the rage.
Well, the antichrist should be in and around this AI thing for one particular reason: The devil cannot create anything of his own because he is not God, by definition. We have already observationally defined generative AI as something that cannot create anything novel in the sense it cannot output anything it has never seen (cannot create new, always a re-assortment of what is ). In that way , AI is a perfect mimicry…
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#108If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
Those 27.3% are still in the ballpark of modern models: - Sonnet 5 - 12.4% - Luna - 17.3% - Grok 4.6 - 20.3% - Sol - 37.3% - GLM 5.3 - 41.8% - Opus 5 - 51.8%
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#109Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked? https://www.youtube.com/watch?v=tNmgmwEtoWE As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in perfor…
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#110Earlier quoted context omitted.
Those 27.3% are still in the ballpark of modern models: - Sonnet 5 - 12.4% - Luna - 17.3% - Grok 4.6 - 20.3% - Sol - 37.3% - GLM 5.3 - 41.8% - Opus 5 - 51.8%
Astra is 58%. The current title says it's "rivaling Astra"