Live data from Hacker News

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

cognition.com

11–20 of 201 posts

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#11

Please correct me if I'm wrong, but this appears to require Devin to use? I'm disappointed to see I need to use a bespoke platform to interact with this agent, to the point that I probably won't be trying it.

SWE-2 is free to use for users like yourself for the next month, and almost all usage should be supported via our CLI (https://docs.devin.ai/cli)

:)

Disclaimer: I work at Cognition, although was not involved in SWE-2

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#12
Where are the model stats? Is this open-weights? If not, why would I use this over DeepSeek Flash 4.1?

I think these competing labs need to realize that no one wants another closed-weight model provider... We aren't even happy with the two we have right now, and their days are entirely numbered. If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper).

The big labs love to release their new model and quantize after the first week. You don't have that problem using dirt cheap API rates on OpenRouter. DS 4.1 flash is also faster than fast mode Astra. OAI's subscription rates are good value, but now these new open-weight models are nearly as cheap on API usage rates. I honestly can't wait for the day we're not beholden to the two big labs anymore. No wonder there's so much fear pumping happening at the moment from Anthropic and their funded NGOs.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#13
post #4

"SWE-2 is post-trained from Kimi K3"

Yeah, I'd expect model performance to be super spiky on SWE work, at least they admit it with the name of the model. It's distilled from an already-distilled model. Maybe still worth it if their "64% cheaper" figure holds.

With the Devin subscription even at the 20$ plan, they offered unlimited SWE 1.7 usage. Wondering if they do the same for SWE 2.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#14
post #13

Earlier quoted context omitted.

Yeah, I'd expect model performance to be super spiky on SWE work, at least they admit it with the name of the model. It's distilled from an already-distilled model. Maybe still worth it if their "64% cheaper" figure holds.

With the Devin subscription even at the 20$ plan, they offered unlimited SWE 1.7 usage. Wondering if they do the same for SWE 2.

SWE-2 is free for all subscribers on the CLI to try out for the next month :)

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#16
At work I setup a cloud worker, where i can spin up as many concurrent agents I want, with unlimited fable 5.1 (thanks employer!!).

I now just work from my phone, and speak into the agents as they run. I dont write code and I dont write documents. I work on very complicated distributed systems. I dont open my laptop most days. Its a legacy brick I carry around.

Some of my coworkers are still doing things by hand, and are working long hours to produce 25% of the output (when considering hours worked). I stay quiet with my setup. We are in the end times for this job for the people that can see clearly how to automate their own job

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#18
post #4

"SWE-2 is post-trained from Kimi K3"

I presume post training is significantly easier than the distillation/training the top Chinese labs are doing.

I wonder if, similar to the American labs, they'll become stingy with their weights once they start getting immediately undercut by a wave of slightly better derived models.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#20

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

Yeah, this echoes my thoughts. I will be very surprised if a model with 2.8T parameters reaches the intelligence and capabilities of 10T parameter models. RL can take things far, but not that far.
Post reply on HN