Live data from Hacker News

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

cognition.com

51–60 of 201 posts

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#51

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.

> Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"?

Yes.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#52
post #28

Earlier quoted context omitted.

Yeah, I'd expect model performance to be super spiky on SWE work, at least they admit it with the name of the model. It's distilled from an already-distilled model. Maybe still worth it if their "64% cheaper" figure holds.

I don't think you know what distill means

I guess I don't. Does post-training from another (larger) model not fall under the umbrella of distillation? I'd imagine it leads to the same spiky-ness issues...?

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#53
post #11

Please correct me if I'm wrong, but this appears to require Devin to use? I'm disappointed to see I need to use a bespoke platform to interact with this agent, to the point that I probably won't be trying it.

SWE-2 is free to use for users like yourself for the next month, and almost all usage should be supported via our CLI ( https://docs.devin.ai/cli ) :) Disclaimer: I work at Cognition, although was not involved in SWE-2

I just gave it a try and it doesn't appear to be free, it used up some of my on demand usage. It does say 75% off though. Seems like for Pro subscribers SWE-1.7 is free, maybe SWE-2 is free for them?

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#54

As an Econ graduate, pretty cool seeing Pareto in the "AI-bro" zeitgeist. Slightly surreal watching a 1906 welfare economics idea get rediscovered as a plotting convention. The original, if anyone fancies 579 pages of Italian: https://archive.org/details/manualedieconomi00pareuoft . There is an English translation somewhere.

Pareto leaked out of the sociology/econ bubble a long time ago :) Pareto principle, Pareto efficiency, Pareto distribution have been in the pop-sci buzzwords for quite awhile, I probably encountered it first in the 4-Hour Workweek. I don't think you can read a self-help book without the author introducing it as a groundbreaking principle to live your life by.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#55
post #40

Earlier quoted context omitted.

A lot of people seem to be the Antichrist these days. I’d have though we’d get a lull after the millennium but it’s all the rage.

Well, the antichrist should be in and around this AI thing for one particular reason: The devil cannot create anything of his own because he is not God, by definition. We have already observationally defined generative AI as something that cannot create anything novel in the sense it cannot output anything it has never seen (cannot create new, always a re-assortment of what is ). In that way , AI is a perfect mimicry…

What’s the antichrists goal though?

Create hell on earth or turn us all into heretics or something else?

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#56

As an Econ graduate, pretty cool seeing Pareto in the "AI-bro" zeitgeist. Slightly surreal watching a 1906 welfare economics idea get rediscovered as a plotting convention. The original, if anyone fancies 579 pages of Italian: https://archive.org/details/manualedieconomi00pareuoft . There is an English translation somewhere.

Pareto frontiers are pretty commonly invoked to describe tradeoffs in computer science and have been for quite a while. I remember the term being used in one of my early algorithms courses to describe the tradeoff between data structures with fast writes, ones with fast reads and ones that tried to balance the two.

IIRC cognition boasted about hiring a lot of competitive programmers and algorithms experts back when they released Devin, so it tracks that they'd use the term.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#58

Earlier quoted context omitted.

that's why after 1 year of product development of these AI 20x maxxed speed, we reached AGI 'wizards', there's really no difference in output, outstanding bugs no longer get solved and sites still suck, even doing things that were just regular development 20 years ago. Are you sure they aren't only producing 2.5% of your output that you manage just by farting into your phone? Are you sure it's 25% really? Seems way t…

please keep thinking this so i can relax with my automated job

It's not you, it's X... but what would you expect of a nepo-baby economy of little swines. This is like the nepo wet-dream on steroids. Incompetence and delulu

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#59
post #11

Please correct me if I'm wrong, but this appears to require Devin to use? I'm disappointed to see I need to use a bespoke platform to interact with this agent, to the point that I probably won't be trying it.

SWE-2 is free to use for users like yourself for the next month, and almost all usage should be supported via our CLI ( https://docs.devin.ai/cli ) :) Disclaimer: I work at Cognition, although was not involved in SWE-2

Your own CLI? Not even a /v1/chat/completions API? Is your business model based on pretending LLMs are not an interchangeable commodity already?

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#60
post #23

Earlier quoted context omitted.

Closed weights AND benchmaxxed. Somehow this company raised 2bil at a 48bil valuation. Pure insanity. I feel bad for their investors (not really, but... Still). Andreessen Horowitz is being played like a fiddle.

The Cursor acquisition shows that it’s possible for these valuations to be justified. But Cursor was more successful and bent the truth much less. While I wouldn’t expect anything good for Cognition’s fate, it’s a much safer bet than Thinking Machines, SSI, and some others. Though they’ll be in big trouble if the more talented Chinese labs stop letting them repackage their work.

The cursor acquisition just shows that there’s always a dumber shithead out there. Though vaporizing Elon Musks money is about as pure of a good as there is out there these days.
Post reply on HN