Live data from Hacker News

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

cognition.com

151–160 of 201 posts

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#151
post #35

Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked? https://www.youtube.com/watch?v=tNmgmwEtoWE As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in perfor…

As others have mentioned, it's really matured a lot and at this point is one of the best cloud-hosted, team-managed coding agents, when factoring overall UX, testing and QA lifecycle via its sandboxes, and its ability to be controlled with an API. We use it quite heavily.

It's coming from a different starting place than Claude Code or Codex are as individually controlled single-developer tools. Devin has been more persistent in pursuing the direction of something that operates more autonomously at the team level, as a peer. And while it might be slightly behind in raw harness ability (maybe?) it's probably ahead on the team-focus.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#152

Earlier quoted context omitted.

OpenAI just paused new subscriptions to their $200 plan. They are in a rock and a hard place. Obviously the Astras and Fables of the world are exponentially more expensive, but for...less than exponential returns. The question is whether they can leverage the marginal advantage into something that justifies the diminishing returns before the bottom catches up to them. On the one hand you, if you bought a lot of compu…

I could buy four of them for ~20k. That's like four years of ChatGPT + Claude subscription. Eight years if only ChatGPT, or sixteen years of the Pro 5x subscription.

It's not particularly good value if you are just comparing $ with no other context. Point is just that it's accessible and so now Anthropic and OpenAI need to make both a performance proposition and a value proposition.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#153
post #151
post #35

Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked? https://www.youtube.com/watch?v=tNmgmwEtoWE As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in perfor…

As others have mentioned, it's really matured a lot and at this point is one of the best cloud-hosted, team-managed coding agents, when factoring overall UX, testing and QA lifecycle via its sandboxes, and its ability to be controlled with an API. We use it quite heavily. It's coming from a different starting place than Claude Code or Codex are as individually controlled single-developer tools. Devin has been more pe…

Your startup is based around AI coworkers, so I expect you are biased towards overvaluing their usefulness.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#154

Earlier quoted context omitted.

What’s the antichrists goal though? Create hell on earth or turn us all into heretics or something else?

Anti Christ goal: Achieve total global dominance and become the object of worship over God, while killing all those who stay faithful to Jesus Christ. Those who stay faithful see Heaven, those who don’t, see the Lake of Fire. It’s the final separation of the wheat from the chaff. As per Revelations. Thank your for allowing me to edify :)

Why would a Christian be trying to stop the Antichrist? Sounds like the faithful have nothing to lose.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#155
post #11

Please correct me if I'm wrong, but this appears to require Devin to use? I'm disappointed to see I need to use a bespoke platform to interact with this agent, to the point that I probably won't be trying it.

SWE-2 is free to use for users like yourself for the next month, and almost all usage should be supported via our CLI ( https://docs.devin.ai/cli ) :) Disclaimer: I work at Cognition, although was not involved in SWE-2

Hey man, how’s the Poke SOC2 audit going? Must be any day now that it’ll be finished, right?

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#156
post #27

Earlier quoted context omitted.

> If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper). Only in the world where the incumbents don't react. Eg if they saw lots of users moving away, they'd drop prices or do something else.

1M cached tokens on deepseek is $0.006, the big labs can't sell anywhere close to this, they have funders expecting returns and huge overhead. btw I've had a ton of fun with the new deepseek today, I was waiting for my OpenAI 5h limit reset and decided to give it some problems for fun, got pretty great results. Tried some harder problems and still got great results. I don't expect it to be Sol class or anything but I…

I did a lot of that kind of work with the older version of DeepSeek before they upped the prices.

For example, it was quite good to get a decent Sashiko review. Sashiko is a Linux kernel review agent with interchangeable LLM driver. It's very good, but it eats tokens like crazy.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#157

Earlier quoted context omitted.

1M cached tokens on deepseek is $0.006, the big labs can't sell anywhere close to this, they have funders expecting returns and huge overhead. btw I've had a ton of fun with the new deepseek today, I was waiting for my OpenAI 5h limit reset and decided to give it some problems for fun, got pretty great results. Tried some harder problems and still got great results. I don't expect it to be Sol class or anything but I…

Yep. 4.1 Flash is good enough for most routine coding things, but it also makes up for a lot of weakness by being so fast (and cheap of course). I'm willing to tolerate babysitting things a lot more if I know I'll get almost instant results.

You could also have eg Sol do the babysitting.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#158
post #27

Earlier quoted context omitted.

> If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper). Only in the world where the incumbents don't react. Eg if they saw lots of users moving away, they'd drop prices or do something else.

OpenAI just paused new subscriptions to their $200 plan. They are in a rock and a hard place. Obviously the Astras and Fables of the world are exponentially more expensive, but for...less than exponential returns. The question is whether they can leverage the marginal advantage into something that justifies the diminishing returns before the bottom catches up to them. On the one hand you, if you bought a lot of compu…

> But the counter to that is that everyone else is becoming more compute efficient so maybe that advantage isn't what people thought it would be.

I'm not sure? If we have techniques to use the hardware even better, that will make the hardware even more valuable, won't it?

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#159
post #151

Earlier quoted context omitted.

As others have mentioned, it's really matured a lot and at this point is one of the best cloud-hosted, team-managed coding agents, when factoring overall UX, testing and QA lifecycle via its sandboxes, and its ability to be controlled with an API. We use it quite heavily. It's coming from a different starting place than Claude Code or Codex are as individually controlled single-developer tools. Devin has been more pe…

Your startup is based around AI coworkers, so I expect you are biased towards overvaluing their usefulness.

Yes, perhaps fair. But my point was somewhat narrow. I wasn't saying that team-managed agents are good to go for all cases and that they're better than individual dev-managed ones. Just that they have gotten better and that of those Devin has some of the better UX.

Our experience might also not be typical because we have built infrastructure around making Devin and similar agents work better. And for the record no ties to Devin/Cognition. Just pay them too much as a customer.

Post reply on HN