Earlier quoted context omitted.
Those 27.3% are still in the ballpark of modern models: - Sonnet 5 - 12.4% - Luna - 17.3% - Grok 4.6 - 20.3% - Sol - 37.3% - GLM 5.3 - 41.8% - Opus 5 - 51.8%
DeepSeek v4.1 Flash 31.2% Source: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash#compa...
Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
111–120 of 201 posts
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#112Not sure it matters when devin is the most consistently shit product I've used. And yes, I tried again, they wasted the money on the billboards.
I have not met a single person/company that uses Devin… does anyone here actually use it?
It's a great product compared to Copilot. It is also the first AI tool I used heavily outside of creating random images or one off questions.
I'm now using all three, Devin, Claude, Codex. I'm finding Claude and Codex to be much better. One of my biggest gripes is that the web client and desktop client for Devin are two completely different harnesses, so the quality of responses varies greatly.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#113(just to be clear, I am a far-left activist who spends most of my time working on funding Social Security Trust Funds (OASI & DI Solvency), which could impact my search results - this was while I was logged in.)
[1]https://www.google.com/search?q=what%27s+cognition+in+ai or https://imgur.com/a/UdxtnGg
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#114Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked? https://www.youtube.com/watch?v=tNmgmwEtoWE As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in perfor…
I'm surprised they can post-train a closed model off of K3. Is that the norm for open-weight licenses?
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#115If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.
It's really, really difficult to avoid it even when you care to stop yourself; and it's not even just a problem in machine learning, it's the standard failure mode of all minds capable of learning, human, animal, artificial.
Even pure genetics has this problem. Viruses and cancers also demonstrate this behaviour, with the bench being evolution's only option: reproductive success.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#116Earlier quoted context omitted.
SWE-2 is free to use for users like yourself for the next month, and almost all usage should be supported via our CLI ( https://docs.devin.ai/cli ) :) Disclaimer: I work at Cognition, although was not involved in SWE-2
But I don't want to use your CLI. I already have my own harnesses and workflows. The friction is too high to "just try out" a new model like this. It would be preferable if I can evaluate it over, say, open router like all the other models and then decide from there if it's worth downloading a bespoke tool chain for only 1 lab's models
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#117If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
Came here to say the exact same thing! People have to stop paying any attention to coding benchmarks that aren't Terminal Bench 4. I noticed I noticed they didn't include Gemini 3.8, which also murders DeepSWE and Terminal Bench 2.0 -- because they are useless benchmarks now! Of course in a couple months TB4 will also be old hat, so TB5 will have to be the new real benchmark.
That then made me realize that they lower the bars of tied scores so on the site it looks like Astra in second place. Weird. Anyway, yes, so many of these composite benchmark sites are irrelevant if they're not trimming the fat and sticking to the most up-to-date variants.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#118Earlier quoted context omitted.
I'm surprised they can post-train a closed model off of K3. Is that the norm for open-weight licenses?
Yes. If they are really open, you can use them as you wish. Some licenses are less open though.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#119If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
Those 27.3% are still in the ballpark of modern models: - Sonnet 5 - 12.4% - Luna - 17.3% - Grok 4.6 - 20.3% - Sol - 37.3% - GLM 5.3 - 41.8% - Opus 5 - 51.8%
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#120> SWE-2 is post-trained from Kimi K3 On the one hand I would have expected a completely new model, on the other hand it's an RL-ed K3 go Fable 5 capabilities, which demonstrate that this is probably possible, which is nice.
why are all American AI models basically Kimi in a trench coat