Live data from Hacker News

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

cognition.com

91–100 of 201 posts

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#91

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

When a benchmark becomes a target, it's no longer a good benchmark...

people are just fighting for numbers.. i don't fundamentally see the model being better

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#92

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

Came here to say the exact same thing! People have to stop paying any attention to coding benchmarks that aren't Terminal Bench 4.

I noticed I noticed they didn't include Gemini 3.8, which also murders DeepSWE and Terminal Bench 2.0 -- because they are useless benchmarks now!

Of course in a couple months TB4 will also be old hat, so TB5 will have to be the new real benchmark.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#93
post #83

Earlier quoted context omitted.

A friend recently pushed me to try out their coding platform, Devin, after I decided to move away from Cursor. I had the same reaction: "What, the con artists from like 2024?" But after some cajoling, I gave it a shot and was pleasantly surprised. I guess they learned their lessons, grew up, and are doing good work now, maybe?

a reminder that money does enable you to make mistakes and buys you the ability to correct from them

> and buys you the ability to correct from them

Or at the very least, make more mistakes.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#94

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.

[deleted]

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#95

Not sure it matters when devin is the most consistently shit product I've used. And yes, I tried again, they wasted the money on the billboards.

I have not met a single person/company that uses Devin… does anyone here actually use it?

I use it, and have been happy with it for the most part. Like sibling, I use it for GPT-5.6-Sol and Opus work, and use their free models (GLM 5.2 for the past few months, trying SWE-2 now).

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#96

Not sure it matters when devin is the most consistently shit product I've used. And yes, I tried again, they wasted the money on the billboards.

I have not met a single person/company that uses Devin… does anyone here actually use it?

my friends at infosys are being trained on it

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#97
post #87

> SWE-2 is post-trained from Kimi K3 On the one hand I would have expected a completely new model, on the other hand it's an RL-ed K3 go Fable 5 capabilities, which demonstrate that this is probably possible, which is nice.

why are all American AI models basically Kimi in a trench coat

No one in the US is going to fund pretty good open source with VC money.

US has OpenAI/ Anthropic/ Google/ meta/ SpaceX atleast trying to make frontier foundation models 2 are using VC money + cash flow, last 3 are mainly cash flow + equity and debt.

China has state banks and similar willing to fund lower margin open source labs.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#98

Earlier quoted context omitted.

please keep thinking this so i can relax with my automated job

It's not you, it's X... but what would you expect of a nepo-baby economy of little swines. This is like the nepo wet-dream on steroids. Incompetence and delulu

[deleted]

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#99
post #11

Earlier quoted context omitted.

SWE-2 is free to use for users like yourself for the next month, and almost all usage should be supported via our CLI ( https://docs.devin.ai/cli ) :) Disclaimer: I work at Cognition, although was not involved in SWE-2

Your own CLI? Not even a /v1/chat/completions API? Is your business model based on pretending LLMs are not an interchangeable commodity already?

they are an agent company not a model provider, is this that difficult to comprehend?

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#100

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.

[deleted]
Post reply on HN