Not sure it matters when devin is the most consistently shit product I've used. And yes, I tried again, they wasted the money on the billboards.
Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
61–70 of 201 posts
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#62Earlier quoted context omitted.
This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.
> Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? A model that was released a couple months ago scores 50% higher than SWE-2, a model released today, on an out-of-sample benchmark. Can I say I’ve come out of this more impressed with Sol? Like you said, TB2 is saturated. Nobody would bat an eyelash at 90%. And yet here comes SWE-2 coming off top rope with an emphatic 92.4%. this is the definition o…
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#63Earlier quoted context omitted.
Well, the antichrist should be in and around this AI thing for one particular reason: The devil cannot create anything of his own because he is not God, by definition. We have already observationally defined generative AI as something that cannot create anything novel in the sense it cannot output anything it has never seen (cannot create new, always a re-assortment of what is ). In that way , AI is a perfect mimicry…
What’s the antichrists goal though? Create hell on earth or turn us all into heretics or something else?
Achieve total global dominance and become the object of worship over God, while killing all those who stay faithful to Jesus Christ.
Those who stay faithful see Heaven, those who don’t, see the Lake of Fire. It’s the final separation of the wheat from the chaff.
As per Revelations. Thank your for allowing me to edify :)
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#64Earlier quoted context omitted.
Is Sol benchmaxxed? Of course it is. Altman was caught in previous attempts trying to game benchmarks, does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with?
> Altman was caught in previous attempts trying to game benchmarks Sounds like something you just made up, or maybe you read it on some other Reddit/HN post and started repeating it because it aligned with your biases. > does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with? I don't think "OpenAI" is equivalent to "Sam Altman." I think if OpenAI was inten…
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#65If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.
Too easy to game the numbers, and too easy to baselessly accuse companies of gaming the numbers, not to mention how you even define that.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#66If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
First, almost all models are within spitting distances of eachother.
Second, it never translates to being better for my own workloads.
You just need to make your own benchmarks.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#67Earlier quoted context omitted.
What’s the antichrists goal though? Create hell on earth or turn us all into heretics or something else?
Anti Christ goal: Achieve total global dominance and become the object of worship over God, while killing all those who stay faithful to Jesus Christ. Those who stay faithful see Heaven, those who don’t, see the Lake of Fire. It’s the final separation of the wheat from the chaff. As per Revelations. Thank your for allowing me to edify :)
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#68Where are the model stats? Is this open-weights? If not, why would I use this over DeepSeek Flash 4.1? I think these competing labs need to realize that no one wants another closed-weight model provider... We aren't even happy with the two we have right now, and their days are entirely numbered. If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the cl…
This really just exists so cognition can stop spending API tokens with Anthropic or OpenAI. Basically any successful AI based service will do this because at scale the frontier models are expensive and you’ll have enough data to fine tune your own. Same reason Harvey is doing models now and basically every other provider
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#69If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#70Not sure it matters when devin is the most consistently shit product I've used. And yes, I tried again, they wasted the money on the billboards.
I have not met a single person/company that uses Devin… does anyone here actually use it?
I used to use windsurf as my main editor until they changed their pricing model. Now i use it just to burn my weekly tokens on fable/astra if i remember to that on a task and that's it.