Live data from Hacker News

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

cognition.com

71–80 of 201 posts

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#71
post #35

Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked? https://www.youtube.com/watch?v=tNmgmwEtoWE As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in perfor…

A friend recently pushed me to try out their coding platform, Devin, after I decided to move away from Cursor. I had the same reaction: "What, the con artists from like 2024?" But after some cajoling, I gave it a shot and was pleasantly surprised. I guess they learned their lessons, grew up, and are doing good work now, maybe?

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#72

Earlier quoted context omitted.

Anti Christ goal: Achieve total global dominance and become the object of worship over God, while killing all those who stay faithful to Jesus Christ. Those who stay faithful see Heaven, those who don’t, see the Lake of Fire. It’s the final separation of the wheat from the chaff. As per Revelations. Thank your for allowing me to edify :)

So, we're rooting for the Anti Christ then? I might have to change my stance on generative "ai" then.

[deleted]

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#73

As an Econ graduate, pretty cool seeing Pareto in the "AI-bro" zeitgeist. Slightly surreal watching a 1906 welfare economics idea get rediscovered as a plotting convention. The original, if anyone fancies 579 pages of Italian: https://archive.org/details/manualedieconomi00pareuoft . There is an English translation somewhere.

Pareto frontiers are pretty commonly invoked to describe tradeoffs in computer science and have been for quite a while. I remember the term being used in one of my early algorithms courses to describe the tradeoff between data structures with fast writes, ones with fast reads and ones that tried to balance the two. IIRC cognition boasted about hiring a lot of competitive programmers and algorithms experts back when t…

It's interesting watching people throw about pareto frontiers sort of like how RF nerds approach the shannon limit (in a practical real world sense of the term, like charting possible modulations/data rates on a two way satellite modem's manufacturer datasheet).

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#74
I think I'm probably in the minority here, but for my line of work, the software engineering and coding is only a small part of the work. I write simulation software, so a deep understanding of physics, math, and how they can be applied to the software is absolutely crucial. I'm assuming this model is tuned to be more focused on SWE topics, and the very reason we seek "multidisciplinary" hires is the also why I actually need a jack-of-all-trades model to back my coding agents.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#75

At work I setup a cloud worker, where i can spin up as many concurrent agents I want, with unlimited fable 5.1 (thanks employer!!). I now just work from my phone, and speak into the agents as they run. I dont write code and I dont write documents. I work on very complicated distributed systems. I dont open my laptop most days. Its a legacy brick I carry around. Some of my coworkers are still doing things by hand, and…

[flagged]

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#78
post #74

I think I'm probably in the minority here, but for my line of work, the software engineering and coding is only a small part of the work. I write simulation software, so a deep understanding of physics, math, and how they can be applied to the software is absolutely crucial. I'm assuming this model is tuned to be more focused on SWE topics, and the very reason we seek "multidisciplinary" hires is the also why I actua…

I'm also in simulation software! Wondering which models you are finding helpful, the models I'm using for general SWE skills are horrible at our simulations and even basic physics/engineering calculation and intuition

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#79

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

Yeah DeepSeek V4.1 beats this by 15 percent (4 points) on terminal bench 4.0

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#80

Earlier quoted context omitted.

The Cursor acquisition shows that it’s possible for these valuations to be justified. But Cursor was more successful and bent the truth much less. While I wouldn’t expect anything good for Cognition’s fate, it’s a much safer bet than Thinking Machines, SSI, and some others. Though they’ll be in big trouble if the more talented Chinese labs stop letting them repackage their work.

The cursor acquisition just shows that there’s always a dumber shithead out there. Though vaporizing Elon Musks money is about as pure of a good as there is out there these days.

Improvements in models and products coming out of SpaceXAI since the acquisition would seem to disagree with you.
Post reply on HN