Live data from Hacker News

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

cognition.com

181–190 of 201 posts

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#181
post #176

If everything basically rivals Fable, then why is everything still using it for comparison?

Have you spent at least 10 seconds thinking about it or are you asking just out of spite?

Let me grab my calculator and add up the time I have spent reading about model releases since Fable has been released. It seems they all place themselves relative to Fable. I'm sure that time has added up to far greater than 10 seconds. At some point, it ceased to be a meaningful differentiation. This is especially true when I put the model through real usage.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#182

Earlier quoted context omitted.

It's not you, it's X... but what would you expect of a nepo-baby economy of little swines. This is like the nepo wet-dream on steroids. Incompetence and delulu

If you are too dumb or too lazy to figure out what this guy has, then you are the problem and will be looked at as a relic. Watching the AI slop my sales reps put in their emails is disgusting but the reply telling them how great of a job they are doing and how insightful their email was says differently. Many people are laughing to the bank while you are still running `--help` to figure out how to run a complex comm…

Maybe it's you that needs to learn how to run `--help` or ask an AI how to not cry about burnout on open source instead? I don't get it, you should just be cruising on auto-pilot now.

The problem is retards that can only function on a cocktail of drugs, and as they were never good at anything other than anal retentive stuff built and continue to build these retarded systems. Those peddling RoR apps even when they couldn't serve more than 3 or 4 concurrent requests, JS backends to handle complex workflows that even after 2 years of dev. still have bugs and accrued a sprawl of crap to hide the issues of their own making, etc, and yet charge thousands of dollars, those that write shit software that's not even worth to clean your ass with, even though they have 20 years of experience, but then go give conferences and write books about their amazing architectural skills, those that write utils behind the "oh, it's open source, if you don't like it just fork it" and due to marketing get their crap everywhere, while making holes everywhere for their paycheques. Or the nepo babies that need their mexico border run to get their fix so they can have these "humanity changing" ideas? I bet they're the same that before would weasel a 2 week sprint to change the borders of a button. Or burn through 10k in meetings for irrelevant crap. Or get VC funding for a CSS styling company or a two prompt company. Or go on about the value of ideas, but then can't even get that going without outsourcing or an AI to help them have those same "ideas".

Ultimately, you just need to turn into a little pig and party in the pigsty, it's not that difficult either, they say pigs are very close anatomically to humans.

At least AI can help untangle the crap the anal retentive retards have built, and thank god, the pig-mor, this society can't even fuck to replacement levels (perhaps they'll manage now with AI).

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#183

Earlier quoted context omitted.

This really just exists so cognition can stop spending API tokens with Anthropic or OpenAI. Basically any successful AI based service will do this because at scale the frontier models are expensive and you’ll have enough data to fine tune your own. Same reason Harvey is doing models now and basically every other provider

Couldn't they just grab and run an open weight model to save on API tokens?

You get better performance if you also finetune it for your task

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#184

SWE-1.5 was surprisingly good when I used it last. I feel like Cognition is one of the solid players that’s flying a bit under the radar while Anthropic and OpenAI race to IPO.

Odd to get downvotes simply for sharing my experience. Like it or not, Cognition has a good frontier model and they are building serious products. They are working hard which is how you become successful. Sorry if that ruffles your feathers.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#185

Not sure it matters when devin is the most consistently shit product I've used. And yes, I tried again, they wasted the money on the billboards.

Well, where VC money is involved, the aim of these ads placed in SF is not exactly to convince any potential users.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#186
post #140

Earlier quoted context omitted.

I just gave it a try and it doesn't appear to be free, it used up some of my on demand usage. It does say 75% off though. Seems like for Pro subscribers SWE-1.7 is free, maybe SWE-2 is free for them?

Did you use it via the CLI or Desktop? It's 75% off in cloud and free to run on your device.

Oh I tried in cloud. I'll give it another shot

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#188

> SWE-2 is post-trained from Kimi K3 On the one hand I would have expected a completely new model, on the other hand it's an RL-ed K3 go Fable 5 capabilities, which demonstrate that this is probably possible, which is nice.

I thought it would be GLM based, as Devin has had free GLM-5.2 for a while now

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#189
post #126

Earlier quoted context omitted.

GLM 5.3 looks strange, because of this Chinese labs benchmaxx moto. So rather they have emergent abilities or... Also a lot of questions to benchmark because opus 5 is completely useless model right now. I think that the main problem with opus that they try to solve context size optimization problem, and that is the main reason why it speaks like alien with only one technical dictionary at hand. So why it is so good?

Opus 5 generates really good code and terminal commands though. It's just bad at the accompanying text it tells you. These benchmarks don't grade the text generation of the response I don't think, only the task outcome.

Opus 5 keeps making howlingly stupid errors, like one recently where a regexp would catch an invalid date because -\d{2} won’t match -00

Seriously.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#190

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.

TB4 has not been saturated yet.

They are all gaming these benchmarks, it is perfectly reasonable not to trust any of them.

Post reply on HN