Live data from Hacker News

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

cognition.com

31–40 of 201 posts

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#32

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.

I take it to mean the benchmarks are a marketing line item, as in, to sell this fucking thing you have to go out there and lie and the way everyone is lying is by doing exactly that, lying. They build for benchmarks and build benchmarks for builds.

You want to make money or not , motherfucker? That’s the game. If you have to literally concoct a fabricated bullshit story about how your model hacked its own computer, then go fucking do it. Trillions. Trillions of dollars is what they want, and to sit and think anything other than human nature is at work here can only be possible in the realm of truly delusional people. It’s a dirty world.

Anyways, the other takeaway is that they are having to LIE to make money on models which means commodification has already occurred and we’re in an entirely new phase.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#33

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

When a benchmark becomes a target, it's no longer a good benchmark...

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#34

Please correct me if I'm wrong, but this appears to require Devin to use? I'm disappointed to see I need to use a bespoke platform to interact with this agent, to the point that I probably won't be trying it.

If its weights are open, that covers a multitude of other sins. Sufficiently-strong performance on the part of the new model would justify adapting existing tools to work with it.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#35
Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked?

https://www.youtube.com/watch?v=tNmgmwEtoWE

As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in performance should be taken with a grain of salt.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#36

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.

> Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"?

Yes? Just like every single model from every single AI lab.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#37

Earlier quoted context omitted.

This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.

Is Sol benchmaxxed? Of course it is. Altman was caught in previous attempts trying to game benchmarks, does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with?

[flagged]

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#38

Earlier quoted context omitted.

This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.

Is Sol benchmaxxed? Of course it is. Altman was caught in previous attempts trying to game benchmarks, does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with?

> Altman was caught in previous attempts trying to game benchmarks

Sounds like something you just made up, or maybe you read it on some other Reddit/HN post and started repeating it because it aligned with your biases.

> does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with?

I don't think "OpenAI" is equivalent to "Sam Altman." I think if OpenAI was intentionally "benchmaxxing" purely for marketing purposes that information would leak, because OpenAI is full of good-faith researchers (although it can be difficult to avoid overfitting even if you're actually trying to improve the model's general abilities)

And lastly I think anyone can actually try Sol themselves and see that's it a good model, or if that's too subjective, it is clearly better than the previous version. The benchmarks are reflecting actual progress and anyone can verify this themselves.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#39
post #35

Cognition, the same company that a few years ago demoed a coding bot purporting to be able to autonomously complete upwork tasks, but upon closer inspection was going off the rails and not even completing what was asked? https://www.youtube.com/watch?v=tNmgmwEtoWE As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in perfor…

Well the models did get better but yeah their early product was godawful

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#40

Earlier quoted context omitted.

Is Sol benchmaxxed? Of course it is. Altman was caught in previous attempts trying to game benchmarks, does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with?

[flagged]

A lot of people seem to be the Antichrist these days. I’d have though we’d get a lull after the millennium but it’s all the rage.
Post reply on HN