Live data from Hacker News

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

cognition.com

191–200 of 201 posts

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#191

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

> benchmaxxed

An aside: When did talking like incels became cool?

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#192

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%). Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"

This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.

> Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"?

Probably

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#193
post #115

Earlier quoted context omitted.

This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.

Benchmaxxing is the default case, and always has been. It's really, really difficult to avoid it even when you care to stop yourself; and it's not even just a problem in machine learning, it's the standard failure mode of all minds capable of learning, human, animal, artificial. Even pure genetics has this problem. Viruses and cancers also demonstrate this behaviour, with the bench being evolution's only option: repr…

Cancers being distinct living beings by itself is controversial, but you unnecessarily anthropomorphizing.

The truth is much more mundane. It's just Goodhart's Law.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#194
post #115

Earlier quoted context omitted.

Benchmaxxing is the default case, and always has been. It's really, really difficult to avoid it even when you care to stop yourself; and it's not even just a problem in machine learning, it's the standard failure mode of all minds capable of learning, human, animal, artificial. Even pure genetics has this problem. Viruses and cancers also demonstrate this behaviour, with the bench being evolution's only option: repr…

You can't justify it being OK just because it's common. Here, this just makes benchmarks into a low signal and useless marketing number once people get numb to all the 99%s. Also no idea what viruses and cancers have to do with this. Cancer is surely very poor reproductively because they never spread to other hosts.

> Also no idea what viruses and cancers have to do with this. Cancer is surely very poor reproductively because they never spread to other hosts

The host's survival is not the benchmark of the reproductive unit, which are the cancer's cells short-term reproduction.

The distinction is the reason benchmaxxing is in fact not good: the benchmark is only approximately related to what people actually care about. You do want your cells to reproduce sucessfully, after all; you just also want some emergency stop buttons for when they go wrong, and those things failing is your body's benchmark rather than your cell's benchmark.

(There's at least two examples of cancers that can be spread from host to host; lupine genital and taxmanian devil nasal, IIRC)

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#195
post #115

Earlier quoted context omitted.

Benchmaxxing is the default case, and always has been. It's really, really difficult to avoid it even when you care to stop yourself; and it's not even just a problem in machine learning, it's the standard failure mode of all minds capable of learning, human, animal, artificial. Even pure genetics has this problem. Viruses and cancers also demonstrate this behaviour, with the bench being evolution's only option: repr…

Cancers being distinct living beings by itself is controversial, but you unnecessarily anthropomorphizing. The truth is much more mundane. It's just Goodhart's Law.

If people already knew about Goodhart's Law, I wouldn't feel it necessay to give illustrated examples.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#196
post #97
post #87

Earlier quoted context omitted.

why are all American AI models basically Kimi in a trench coat

No one in the US is going to fund pretty good open source with VC money. US has OpenAI/ Anthropic/ Google/ meta/ SpaceX atleast trying to make frontier foundation models 2 are using VC money + cash flow, last 3 are mainly cash flow + equity and debt. China has state banks and similar willing to fund lower margin open source labs.

I'd be pretty surprised if someone told me few years ago communist China, state banks would become the main founders of open source compute and our last hope against monopolists like Musk, the whole bunch at OpenAI and so on.

To be fair Zuck is releasing fairly capable open source models, but Chinese labs are way ahead.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#197
post #111

Earlier quoted context omitted.

DeepSeek v4.1 Flash 31.2% Source: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash#compa...

I have been using it today the whole day and it is definitely better than Luna. Great that bench agrees.

Can you tell something about your tasks? I am pondering both models for cost saving.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#199
post #40

Earlier quoted context omitted.

A lot of people seem to be the Antichrist these days. I’d have though we’d get a lull after the millennium but it’s all the rage.

Well, the antichrist should be in and around this AI thing for one particular reason: The devil cannot create anything of his own because he is not God, by definition. We have already observationally defined generative AI as something that cannot create anything novel in the sense it cannot output anything it has never seen (cannot create new, always a re-assortment of what is ). In that way , AI is a perfect mimicry…

Oh goodness. This is on HN?

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#200
post #149
post #65

Earlier quoted context omitted.

Fixed benchmarks will be debunked eventually I'd think. Better to use synthetic problems. Too easy to game the numbers, and too easy to baselessly accuse companies of gaming the numbers, not to mention how you even define that.

Wouldn't work. A dynamic but verifiable problem. Is just a perfect target for a RL environment. If you don't have the verifiable part the benchmark is useless, or really expensive with human review. (Or just open ended)

There doesn't need to be a single correct response. Generate "novel" problems with a set of acceptable solutions/outcomes, then verify that the response satisfies the criteria.
Post reply on HN