Live data from Hacker News

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

cognition.com

141–150 of 201 posts

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#141

Not sure it matters when devin is the most consistently shit product I've used. And yes, I tried again, they wasted the money on the billboards.

> And yes, I tried again

This made me laugh a bit. I was forced to do an evaluation of their shit product twice due to being backed by the same PE firm; "take a look at it again, it's much better now". It sucked the second time also...

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#142
post #35

Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked? https://www.youtube.com/watch?v=tNmgmwEtoWE As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in perfor…

A friend recently pushed me to try out their coding platform, Devin, after I decided to move away from Cursor. I had the same reaction: "What, the con artists from like 2024?" But after some cajoling, I gave it a shot and was pleasantly surprised. I guess they learned their lessons, grew up, and are doing good work now, maybe?

Devin lacks many features and doesn't even have an exec mode. Other than they subsidize the subscription, the added value here is low.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#143

Earlier quoted context omitted.

A friend recently pushed me to try out their coding platform, Devin, after I decided to move away from Cursor. I had the same reaction: "What, the con artists from like 2024?" But after some cajoling, I gave it a shot and was pleasantly surprised. I guess they learned their lessons, grew up, and are doing good work now, maybe?

Devin CLI is easily the worst-in-class coding harness I've ever been subjected to using, rife with bugs (up to and including dropping answers to the question tool), so I can't say I'm exactly inspired to try anything else the company produces.

Had the same experience on it when it was going by Windsurf. Had to wire up a skill hooked to terminal runs or else it would hang or freeze and never finish literally every single time

I see issues with other harnesses too but not with the regularity I was getting from this. And the one moat they had with the better UI for per-project multi-agent orch disappeared and now is standardized

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#144

Earlier quoted context omitted.

This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.

> Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? A model that was released a couple months ago scores 50% higher than SWE-2, a model released today, on an out-of-sample benchmark. Can I say I’ve come out of this more impressed with Sol? Like you said, TB2 is saturated. Nobody would bat an eyelash at 90%. And yet here comes SWE-2 coming off top rope with an emphatic 92.4%. this is the definition o…

> A model that was released a couple months ago scores 50% higher than SWE-2, a model released today, on an out-of-sample benchmark. Can I say I’ve come out of this more impressed with Sol?

Wait a second, are we taking into account the massive difference in terms of resources of these two companies?

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#146

Earlier quoted context omitted.

> Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? A model that was released a couple months ago scores 50% higher than SWE-2, a model released today, on an out-of-sample benchmark. Can I say I’ve come out of this more impressed with Sol? Like you said, TB2 is saturated. Nobody would bat an eyelash at 90%. And yet here comes SWE-2 coming off top rope with an emphatic 92.4%. this is the definition o…

> A model that was released a couple months ago scores 50% higher than SWE-2, a model released today, on an out-of-sample benchmark. Can I say I’ve come out of this more impressed with Sol? Wait a second, are we taking into account the massive difference in terms of resources of these two companies?

Irrelevant when they say they're competitive with Fable and Astra. They don't get to then roll that back and then say "but we have less compute!"

You're either competitive or not.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#147
post #78
post #74

I think I'm probably in the minority here, but for my line of work, the software engineering and coding is only a small part of the work. I write simulation software, so a deep understanding of physics, math, and how they can be applied to the software is absolutely crucial. I'm assuming this model is tuned to be more focused on SWE topics, and the very reason we seek "multidisciplinary" hires is the also why I actua…

I'm also in simulation software! Wondering which models you are finding helpful, the models I'm using for general SWE skills are horrible at our simulations and even basic physics/engineering calculation and intuition

I'm in the same field and I find GPT to be better at understanding physics conceptually but Claude is better at writing numerical code. I use cursor so many of my sessions start in GPT and switch to Claude.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#148

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.

I mean, still benchmaxxed, I happen to not consider that a problem

These firms are literally hiring professionals from all fields to teach procedure

To teach processes that can subsequently be done agentically or in automated chains

Its basically infinite permutations of tool calling, except the tools aren't external, they’re baked in upon birth

So yeah still makes sense that the new benchmark has a low score and the older one has a high score. And sure, one day we wont have to debate it and a new model will ace everything. Do you actually want that day to be today?

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#149
post #65

Earlier quoted context omitted.

This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.

Fixed benchmarks will be debunked eventually I'd think. Better to use synthetic problems. Too easy to game the numbers, and too easy to baselessly accuse companies of gaming the numbers, not to mention how you even define that.

Wouldn't work. A dynamic but verifiable problem. Is just a perfect target for a RL environment. If you don't have the verifiable part the benchmark is useless, or really expensive with human review. (Or just open ended)
Post reply on HN