Earlier quoted context omitted.
OpenAI just paused new subscriptions to their $200 plan. They are in a rock and a hard place. Obviously the Astras and Fables of the world are exponentially more expensive, but for...less than exponential returns. The question is whether they can leverage the marginal advantage into something that justifies the diminishing returns before the bottom catches up to them. On the one hand you, if you bought a lot of compu…
> But the counter to that is that everyone else is becoming more compute efficient so maybe that advantage isn't what people thought it would be. I'm not sure? If we have techniques to use the hardware even better, that will make the hardware even more valuable, won't it?
Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
171–180 of 201 posts
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#172Earlier quoted context omitted.
Those 27.3% are still in the ballpark of modern models: - Sonnet 5 - 12.4% - Luna - 17.3% - Grok 4.6 - 20.3% - Sol - 37.3% - GLM 5.3 - 41.8% - Opus 5 - 51.8%
GLM 5.3 looks strange, because of this Chinese labs benchmaxx moto. So rather they have emergent abilities or... Also a lot of questions to benchmark because opus 5 is completely useless model right now. I think that the main problem with opus that they try to solve context size optimization problem, and that is the main reason why it speaks like alien with only one technical dictionary at hand. So why it is so good?
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#173Where are the model stats? Is this open-weights? If not, why would I use this over DeepSeek Flash 4.1? I think these competing labs need to realize that no one wants another closed-weight model provider... We aren't even happy with the two we have right now, and their days are entirely numbered. If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the cl…
This really just exists so cognition can stop spending API tokens with Anthropic or OpenAI. Basically any successful AI based service will do this because at scale the frontier models are expensive and you’ll have enough data to fine tune your own. Same reason Harvey is doing models now and basically every other provider
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#174Earlier quoted context omitted.
This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.
Benchmaxxing is the default case, and always has been. It's really, really difficult to avoid it even when you care to stop yourself; and it's not even just a problem in machine learning, it's the standard failure mode of all minds capable of learning, human, animal, artificial. Even pure genetics has this problem. Viruses and cancers also demonstrate this behaviour, with the bench being evolution's only option: repr…
Also no idea what viruses and cancers have to do with this. Cancer is surely very poor reproductively because they never spread to other hosts.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#175Earlier quoted context omitted.
> Altman was caught in previous attempts trying to game benchmarks Sounds like something you just made up, or maybe you read it on some other Reddit/HN post and started repeating it because it aligned with your biases. > does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with? I don't think "OpenAI" is equivalent to "Sam Altman." I think if OpenAI was inten…
There's this whole discussion going on about agents being more independent now. They don't follow instructions so well, they continue until the problem is done (sometimes too long), they don't ask the user for feedback. That is a kind of benchmaxing: they are made to complete benchmarks tasks and one-offs well, and no longer work well in tandem with the user. Regardless what you call it, it's a divergence between wha…
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#176Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#177If everything basically rivals Fable, then why is everything still using it for comparison?
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#178Earlier quoted context omitted.
Benchmaxxing is the default case, and always has been. It's really, really difficult to avoid it even when you care to stop yourself; and it's not even just a problem in machine learning, it's the standard failure mode of all minds capable of learning, human, animal, artificial. Even pure genetics has this problem. Viruses and cancers also demonstrate this behaviour, with the bench being evolution's only option: repr…
You can't justify it being OK just because it's common. Here, this just makes benchmarks into a low signal and useless marketing number once people get numb to all the 99%s. Also no idea what viruses and cancers have to do with this. Cancer is surely very poor reproductively because they never spread to other hosts.
Cells have a certain optimised DNA mutation rate kept in check by various machinery. Multi cellular life expects each of these little replication machine to co-operate in the grand scheme of running a body. But it's also required in the grand scheme for DNA to mutate a little bit to ensure population variance. So you could say that cancer is the tax paid for having cooperative yet flexible and adaptive nano machinery.
So yes the propensity for cancer developed under evolutioniary pressure towards a non zero level.
The population could have optimised for zero cancer but it would not have paid for itself in terms of overall population adaptability and survival.
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#179Earlier quoted context omitted.
> Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"? A model that was released a couple months ago scores 50% higher than SWE-2, a model released today, on an out-of-sample benchmark. Can I say I’ve come out of this more impressed with Sol? Like you said, TB2 is saturated. Nobody would bat an eyelash at 90%. And yet here comes SWE-2 coming off top rope with an emphatic 92.4%. this is the definition o…
> A model that was released a couple months ago scores 50% higher than SWE-2, a model released today, on an out-of-sample benchmark. Can I say I’ve come out of this more impressed with Sol? Wait a second, are we taking into account the massive difference in terms of resources of these two companies?
Re: Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
#180Earlier quoted context omitted.
There's this whole discussion going on about agents being more independent now. They don't follow instructions so well, they continue until the problem is done (sometimes too long), they don't ask the user for feedback. That is a kind of benchmaxing: they are made to complete benchmarks tasks and one-offs well, and no longer work well in tandem with the user. Regardless what you call it, it's a divergence between wha…
Sounds more like profitmaxxxing to me to be honest.