Live data from Hacker News

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

cognition.com

81–90 of 201 posts

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#81

Earlier quoted context omitted.

Is Sol benchmaxxed? Of course it is. Altman was caught in previous attempts trying to game benchmarks, does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with?

> Altman was caught in previous attempts trying to game benchmarks Sounds like something you just made up, or maybe you read it on some other Reddit/HN post and started repeating it because it aligned with your biases. > does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with? I don't think "OpenAI" is equivalent to "Sam Altman." I think if OpenAI was inten…

There's this whole discussion going on about agents being more independent now. They don't follow instructions so well, they continue until the problem is done (sometimes too long), they don't ask the user for feedback.

That is a kind of benchmaxing: they are made to complete benchmarks tasks and one-offs well, and no longer work well in tandem with the user.

Regardless what you call it, it's a divergence between what the power user wants and what the model developers want, I think.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#82

Earlier quoted context omitted.

What’s the antichrists goal though? Create hell on earth or turn us all into heretics or something else?

Anti Christ goal: Achieve total global dominance and become the object of worship over God, while killing all those who stay faithful to Jesus Christ. Those who stay faithful see Heaven, those who don’t, see the Lake of Fire. It’s the final separation of the wheat from the chaff. As per Revelations. Thank your for allowing me to edify :)

It's actually pretty cool of these old stories how well that maps to a dangerous tyrant.

The roman empire perfectly matched that and most powerful men seem to go that path.

It would also be kinda easy to argue many moderns countries are going down that path.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#83
post #35

Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked? https://www.youtube.com/watch?v=tNmgmwEtoWE As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in perfor…

A friend recently pushed me to try out their coding platform, Devin, after I decided to move away from Cursor. I had the same reaction: "What, the con artists from like 2024?" But after some cajoling, I gave it a shot and was pleasantly surprised. I guess they learned their lessons, grew up, and are doing good work now, maybe?

a reminder that money does enable you to make mistakes and buys you the ability to correct from them

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#85
post #28

Earlier quoted context omitted.

I don't think you know what distill means

I guess I don't. Does post-training from another (larger) model not fall under the umbrella of distillation? I'd imagine it leads to the same spiky-ness issues...?

Distilling you don't have the actual model weights of the teacher. All you have are the teachers answers to a lot of questions. You then teach your own smaller model to answer more similarly to the big teacher model.

Fine tuning you have the actual model weights of the original model, you then train that model to answer in a different (or better) way.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#86

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.

"Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"

Yes! extremely sharp RL-fried model. byte perfect hash gates and soak and smoke tests abound.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#87

> SWE-2 is post-trained from Kimi K3 On the one hand I would have expected a completely new model, on the other hand it's an RL-ed K3 go Fable 5 capabilities, which demonstrate that this is probably possible, which is nice.

why are all American AI models basically Kimi in a trench coat

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#88

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

Those 27.3% are still in the ballpark of modern models: - Sonnet 5 - 12.4% - Luna - 17.3% - Grok 4.6 - 20.3% - Sol - 37.3% - GLM 5.3 - 41.8% - Opus 5 - 51.8%

DeepSeek v4.1 Flash 31.2%

Source: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash#compa...

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#89
post #11

Please correct me if I'm wrong, but this appears to require Devin to use? I'm disappointed to see I need to use a bespoke platform to interact with this agent, to the point that I probably won't be trying it.

SWE-2 is free to use for users like yourself for the next month, and almost all usage should be supported via our CLI ( https://docs.devin.ai/cli ) :) Disclaimer: I work at Cognition, although was not involved in SWE-2

Heads up: it doesn't appear to be available on the Devin CLI (for me as a free user).

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#90
post #35

Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked? https://www.youtube.com/watch?v=tNmgmwEtoWE As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in perfor…

Previous versions were based on Kimi too. I'd consider if it I could access the model outside Devin. No lock in for me, thank you very much.
Post reply on HN