If everything basically rivals Fable, then why is everything still using it for comparison?
Have you spent at least 10 seconds thinking about it or are you asking just out of spite?
Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
181–190 of 201 posts
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#182Earlier quoted context omitted.
It's not you, it's X... but what would you expect of a nepo-baby economy of little swines. This is like the nepo wet-dream on steroids. Incompetence and delulu
If you are too dumb or too lazy to figure out what this guy has, then you are the problem and will be looked at as a relic. Watching the AI slop my sales reps put in their emails is disgusting but the reply telling them how great of a job they are doing and how insightful their email was says differently. Many people are laughing to the bank while you are still running `--help` to figure out how to run a complex comm…
The problem is retards that can only function on a cocktail of drugs, and as they were never good at anything other than anal retentive stuff built and continue to build these retarded systems. Those peddling RoR apps even when they couldn't serve more than 3 or 4 concurrent requests, JS backends to handle complex workflows that even after 2 years of dev. still have bugs and accrued a sprawl of crap to hide the issues of their own making, etc, and yet charge thousands of dollars, those that write shit software that's not even worth to clean your ass with, even though they have 20 years of experience, but then go give conferences and write books about their amazing architectural skills, those that write utils behind the "oh, it's open source, if you don't like it just fork it" and due to marketing get their crap everywhere, while making holes everywhere for their paycheques. Or the nepo babies that need their mexico border run to get their fix so they can have these "humanity changing" ideas? I bet they're the same that before would weasel a 2 week sprint to change the borders of a button. Or burn through 10k in meetings for irrelevant crap. Or get VC funding for a CSS styling company or a two prompt company. Or go on about the value of ideas, but then can't even get that going without outsourcing or an AI to help them have those same "ideas".
Ultimately, you just need to turn into a little pig and party in the pigsty, it's not that difficult either, they say pigs are very close anatomically to humans.
At least AI can help untangle the crap the anal retentive retards have built, and thank god, the pig-mor, this society can't even fuck to replacement levels (perhaps they'll manage now with AI).
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#183Earlier quoted context omitted.
This really just exists so cognition can stop spending API tokens with Anthropic or OpenAI. Basically any successful AI based service will do this because at scale the frontier models are expensive and you’ll have enough data to fine tune your own. Same reason Harvey is doing models now and basically every other provider
Couldn't they just grab and run an open weight model to save on API tokens?
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#184SWE-1.5 was surprisingly good when I used it last. I feel like Cognition is one of the solid players that’s flying a bit under the radar while Anthropic and OpenAI race to IPO.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#185Not sure it matters when devin is the most consistently shit product I've used. And yes, I tried again, they wasted the money on the billboards.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#186Earlier quoted context omitted.
I just gave it a try and it doesn't appear to be free, it used up some of my on demand usage. It does say 75% off though. Seems like for Pro subscribers SWE-1.7 is free, maybe SWE-2 is free for them?
Did you use it via the CLI or Desktop? It's 75% off in cloud and free to run on your device.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#187Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#188> SWE-2 is post-trained from Kimi K3 On the one hand I would have expected a completely new model, on the other hand it's an RL-ed K3 go Fable 5 capabilities, which demonstrate that this is probably possible, which is nice.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#189Earlier quoted context omitted.
GLM 5.3 looks strange, because of this Chinese labs benchmaxx moto. So rather they have emergent abilities or... Also a lot of questions to benchmark because opus 5 is completely useless model right now. I think that the main problem with opus that they try to solve context size optimization problem, and that is the main reason why it speaks like alien with only one technical dictionary at hand. So why it is so good?
Opus 5 generates really good code and terminal commands though. It's just bad at the accompanying text it tells you. These benchmarks don't grade the text generation of the response I don't think, only the task outcome.
Seriously.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#190If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.
They are all gaming these benchmarks, it is perfectly reasonable not to trust any of them.