Live data from Hacker News

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

cognition.com

131–140 of 201 posts

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#131
post #35

Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked? https://www.youtube.com/watch?v=tNmgmwEtoWE As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in perfor…

A friend recently pushed me to try out their coding platform, Devin, after I decided to move away from Cursor. I had the same reaction: "What, the con artists from like 2024?" But after some cajoling, I gave it a shot and was pleasantly surprised. I guess they learned their lessons, grew up, and are doing good work now, maybe?

Their ads in SF are pretty funny. “Remember Devin? It’s good now”. Okay, very self-aware, Cognition.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#132

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

Those 27.3% are still in the ballpark of modern models: - Sonnet 5 - 12.4% - Luna - 17.3% - Grok 4.6 - 20.3% - Sol - 37.3% - GLM 5.3 - 41.8% - Opus 5 - 51.8%

So, better than Sonnet and Luna? lol

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#134
post #78
post #74

I think I'm probably in the minority here, but for my line of work, the software engineering and coding is only a small part of the work. I write simulation software, so a deep understanding of physics, math, and how they can be applied to the software is absolutely crucial. I'm assuming this model is tuned to be more focused on SWE topics, and the very reason we seek "multidisciplinary" hires is the also why I actua…

I'm also in simulation software! Wondering which models you are finding helpful, the models I'm using for general SWE skills are horrible at our simulations and even basic physics/engineering calculation and intuition

I use Claude Opus 4.8 almost exclusively. I have had fairly good experiences with it. One time it derived an entirely novel simulation method different than anything in literature by combining its knowledge about how problems in other fields with similar underlying mathematical structure are solved. That was a bit of a Jacobian Conjecture moment for me.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#135
post #115

Earlier quoted context omitted.

Benchmaxxing is the default case, and always has been. It's really, really difficult to avoid it even when you care to stop yourself; and it's not even just a problem in machine learning, it's the standard failure mode of all minds capable of learning, human, animal, artificial. Even pure genetics has this problem. Viruses and cancers also demonstrate this behaviour, with the bench being evolution's only option: repr…

Why would I adapt myself to unseen environments unnecessarily?

Propagation

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#136
post #115

Earlier quoted context omitted.

Benchmaxxing is the default case, and always has been. It's really, really difficult to avoid it even when you care to stop yourself; and it's not even just a problem in machine learning, it's the standard failure mode of all minds capable of learning, human, animal, artificial. Even pure genetics has this problem. Viruses and cancers also demonstrate this behaviour, with the bench being evolution's only option: repr…

Why would I adapt myself to unseen environments unnecessarily?

Because then you could sit under the palm tree and enjoy your free bananas and relax.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#137
post #115

Earlier quoted context omitted.

Benchmaxxing is the default case, and always has been. It's really, really difficult to avoid it even when you care to stop yourself; and it's not even just a problem in machine learning, it's the standard failure mode of all minds capable of learning, human, animal, artificial. Even pure genetics has this problem. Viruses and cancers also demonstrate this behaviour, with the bench being evolution's only option: repr…

Why would I adapt myself to unseen environments unnecessarily?

You can see possible existing environments even if you explicitly blind yourself to them, through reflections off of environments that you do not blind yourself to. Even with the benchmark excluded, the social zeitgeist that has considered the benchmark and included it or ideas from the benchmark either implicitly or explicitly in their code, documentation, et cetera is still part of your training data.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#138
post #35

Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked? https://www.youtube.com/watch?v=tNmgmwEtoWE As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in perfor…

A friend recently pushed me to try out their coding platform, Devin, after I decided to move away from Cursor. I had the same reaction: "What, the con artists from like 2024?" But after some cajoling, I gave it a shot and was pleasantly surprised. I guess they learned their lessons, grew up, and are doing good work now, maybe?

Devin CLI is easily the worst-in-class coding harness I've ever been subjected to using, rife with bugs (up to and including dropping answers to the question tool), so I can't say I'm exactly inspired to try anything else the company produces.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#139

Not sure it matters when devin is the most consistently shit product I've used. And yes, I tried again, they wasted the money on the billboards.

I have not met a single person/company that uses Devin… does anyone here actually use it?

Unfortunately yes. It's awful, other than that they at least support both open-weight and proprietary models to route to. So I guess their model router is "fine", but the harness? As I said in a sibling thread on this page:

> Devin CLI is easily the worst-in-class coding harness I've ever been subjected to using, rife with bugs (up to and including dropping answers to the question tool), so I can't say I'm exactly inspired to try anything else the company produces.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#140
post #11

Earlier quoted context omitted.

SWE-2 is free to use for users like yourself for the next month, and almost all usage should be supported via our CLI ( https://docs.devin.ai/cli ) :) Disclaimer: I work at Cognition, although was not involved in SWE-2

I just gave it a try and it doesn't appear to be free, it used up some of my on demand usage. It does say 75% off though. Seems like for Pro subscribers SWE-1.7 is free, maybe SWE-2 is free for them?

Did you use it via the CLI or Desktop? It's 75% off in cloud and free to run on your device.
Post reply on HN