Live data from Hacker News

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

cognition.com

171–180 of 201 posts

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#171
post #158

Earlier quoted context omitted.

OpenAI just paused new subscriptions to their $200 plan. They are in a rock and a hard place. Obviously the Astras and Fables of the world are exponentially more expensive, but for...less than exponential returns. The question is whether they can leverage the marginal advantage into something that justifies the diminishing returns before the bottom catches up to them. On the one hand you, if you bought a lot of compu…

> But the counter to that is that everyone else is becoming more compute efficient so maybe that advantage isn't what people thought it would be. I'm not sure? If we have techniques to use the hardware even better, that will make the hardware even more valuable, won't it?

But they aren't really competitive for cost in that size class.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#172
post #126

Earlier quoted context omitted.

Those 27.3% are still in the ballpark of modern models: - Sonnet 5 - 12.4% - Luna - 17.3% - Grok 4.6 - 20.3% - Sol - 37.3% - GLM 5.3 - 41.8% - Opus 5 - 51.8%

GLM 5.3 looks strange, because of this Chinese labs benchmaxx moto. So rather they have emergent abilities or... Also a lot of questions to benchmark because opus 5 is completely useless model right now. I think that the main problem with opus that they try to solve context size optimization problem, and that is the main reason why it speaks like alien with only one technical dictionary at hand. So why it is so good?

Opus 5 generates really good code and terminal commands though. It's just bad at the accompanying text it tells you. These benchmarks don't grade the text generation of the response I don't think, only the task outcome.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#173
post #12

Where are the model stats? Is this open-weights? If not, why would I use this over DeepSeek Flash 4.1? I think these competing labs need to realize that no one wants another closed-weight model provider... We aren't even happy with the two we have right now, and their days are entirely numbered. If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the cl…

This really just exists so cognition can stop spending API tokens with Anthropic or OpenAI. Basically any successful AI based service will do this because at scale the frontier models are expensive and you’ll have enough data to fine tune your own. Same reason Harvey is doing models now and basically every other provider

Couldn't they just grab and run an open weight model to save on API tokens?

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#174
post #115

Earlier quoted context omitted.

This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.

Benchmaxxing is the default case, and always has been. It's really, really difficult to avoid it even when you care to stop yourself; and it's not even just a problem in machine learning, it's the standard failure mode of all minds capable of learning, human, animal, artificial. Even pure genetics has this problem. Viruses and cancers also demonstrate this behaviour, with the bench being evolution's only option: repr…

You can't justify it being OK just because it's common. Here, this just makes benchmarks into a low signal and useless marketing number once people get numb to all the 99%s.

Also no idea what viruses and cancers have to do with this. Cancer is surely very poor reproductively because they never spread to other hosts.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#175
post #81

Earlier quoted context omitted.

> Altman was caught in previous attempts trying to game benchmarks Sounds like something you just made up, or maybe you read it on some other Reddit/HN post and started repeating it because it aligned with your biases. > does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with? I don't think "OpenAI" is equivalent to "Sam Altman." I think if OpenAI was inten…

There's this whole discussion going on about agents being more independent now. They don't follow instructions so well, they continue until the problem is done (sometimes too long), they don't ask the user for feedback. That is a kind of benchmaxing: they are made to complete benchmarks tasks and one-offs well, and no longer work well in tandem with the user. Regardless what you call it, it's a divergence between wha…

Sounds more like profitmaxxxing to me to be honest.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#178
post #115

Earlier quoted context omitted.

Benchmaxxing is the default case, and always has been. It's really, really difficult to avoid it even when you care to stop yourself; and it's not even just a problem in machine learning, it's the standard failure mode of all minds capable of learning, human, animal, artificial. Even pure genetics has this problem. Viruses and cancers also demonstrate this behaviour, with the bench being evolution's only option: repr…

You can't justify it being OK just because it's common. Here, this just makes benchmarks into a low signal and useless marketing number once people get numb to all the 99%s. Also no idea what viruses and cancers have to do with this. Cancer is surely very poor reproductively because they never spread to other hosts.

Cancer is simply cells breaking free of the cooperative jail. Essentially the grey goo scenario of nano machines. Instead of cooperation they just do their own thing.

Cells have a certain optimised DNA mutation rate kept in check by various machinery. Multi cellular life expects each of these little replication machine to co-operate in the grand scheme of running a body. But it's also required in the grand scheme for DNA to mutate a little bit to ensure population variance. So you could say that cancer is the tax paid for having cooperative yet flexible and adaptive nano machinery.

So yes the propensity for cancer developed under evolutioniary pressure towards a non zero level.

The population could have optimised for zero cancer but it would not have paid for itself in terms of overall population adaptability and survival.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#179

Earlier quoted context omitted.

> Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? A model that was released a couple months ago scores 50% higher than SWE-2, a model released today, on an out-of-sample benchmark. Can I say I’ve come out of this more impressed with Sol? Like you said, TB2 is saturated. Nobody would bat an eyelash at 90%. And yet here comes SWE-2 coming off top rope with an emphatic 92.4%. this is the definition o…

> A model that was released a couple months ago scores 50% higher than SWE-2, a model released today, on an out-of-sample benchmark. Can I say I’ve come out of this more impressed with Sol? Wait a second, are we taking into account the massive difference in terms of resources of these two companies?

You should use my model then, I spent about 30$ in electricity and used my existing RTX4090. It is not very good, but can you compare it with others really? You can use this service via a private API with a VPN, email me your credit card details for access.

Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

#180
post #175
post #81

Earlier quoted context omitted.

There's this whole discussion going on about agents being more independent now. They don't follow instructions so well, they continue until the problem is done (sometimes too long), they don't ask the user for feedback. That is a kind of benchmaxing: they are made to complete benchmarks tasks and one-offs well, and no longer work well in tandem with the user. Regardless what you call it, it's a divergence between wha…

Sounds more like profitmaxxxing to me to be honest.

They have plausible deniability on that one: not making any profit.
Post reply on HN