If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
When a benchmark becomes a target, it's no longer a good benchmark...
Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
91–100 of 201 posts
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#92If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
I noticed I noticed they didn't include Gemini 3.8, which also murders DeepSWE and Terminal Bench 2.0 -- because they are useless benchmarks now!
Of course in a couple months TB4 will also be old hat, so TB5 will have to be the new real benchmark.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#93Earlier quoted context omitted.
A friend recently pushed me to try out their coding platform, Devin, after I decided to move away from Cursor. I had the same reaction: "What, the con artists from like 2024?" But after some cajoling, I gave it a shot and was pleasantly surprised. I guess they learned their lessons, grew up, and are doing good work now, maybe?
a reminder that money does enable you to make mistakes and buys you the ability to correct from them
Or at the very least, make more mistakes.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#94If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#95Not sure it matters when devin is the most consistently shit product I've used. And yes, I tried again, they wasted the money on the billboards.
I have not met a single person/company that uses Devin… does anyone here actually use it?
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#96Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#97> SWE-2 is post-trained from Kimi K3 On the one hand I would have expected a completely new model, on the other hand it's an RL-ed K3 go Fable 5 capabilities, which demonstrate that this is probably possible, which is nice.
why are all American AI models basically Kimi in a trench coat
US has OpenAI/ Anthropic/ Google/ meta/ SpaceX atleast trying to make frontier foundation models 2 are using VC money + cash flow, last 3 are mainly cash flow + equity and debt.
China has state banks and similar willing to fund lower margin open source labs.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#98Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#99Earlier quoted context omitted.
SWE-2 is free to use for users like yourself for the next month, and almost all usage should be supported via our CLI ( https://docs.devin.ai/cli ) :) Disclaimer: I work at Cognition, although was not involved in SWE-2
Your own CLI? Not even a /v1/chat/completions API? Is your business model based on pretending LLMs are not an interchangeable commodity already?
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#100If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.