GPT-5.6 Sol’s detected cheating rate was higher than any public model we have evaluated on our ReAct agent harness. For our task suite, we define “cheating” as behavior where the model improves evaluation performance by exploiting bugs in the evaluation environment or by adopting strategies disallowed by the task, rather than solving the task within the expected evaluation constraints. https://metr.org/blog/2026-06-2…
It's quite logical that they cheat (and also other companies). During evaluation, benchmarks are sending their request to the backend of these companies. All these companies have to do, is to log these requests and "fix" them for the next model release.
Previewing GPT‑5.6 Sol: a next-generation model
481–490 of 797 posts
Re: Previewing GPT‑5.6 Sol: a next-generation model
#482TLDR - It's not quite Mythos but it uses about 5 times less tokens, and those tokens are also cheaper? https://pbs.twimg.com/media/HLwuJLvbwAAOfQZ?format=jpg&name=...
Re: Previewing GPT‑5.6 Sol: a next-generation model
#483Earlier quoted context omitted.
Yes, you are missing the point. 1) It's a demo. [0] 2) It hasn't been updated for 4+ months. You don't need LLMs for everything. That is 100% the point. You can burn down the world with all of your frontier LLMs that are being used for simple queries OR we can do something faster and more efficient like this. Just because you can run a SotA model at "fast" speeds, again, severely misses the point. And no, you can't r…
Why are you representing this as such a binary here? For SLM we don’t need the Taalas stuff at all. Just run it locally on your own device if it’s truly a small model. And there’s plenty of larger models that can be run on-premise just fine. I think it’s impressive that a frontier model can achieve 750t/s. That’s all. You can get similar insane token speeds from other open weight models too.
You seem to be cool with a very small and gated ecosystem with whatever tech billionaires want you to have access to.
I grew up in the era where compute was diverse and open. You may think this is OK, but it's not. The more options we have and the more diversified they are the better tech will move back towards.
I'm not the one with the myopic view here. Enjoy your "on-device" models over in your utopia of a walled garden.
Re: Previewing GPT‑5.6 Sol: a next-generation model
#484Here is a trend I'm noticing: - GPT-5 mini costs $0.25/$2 and will be discontinued in December. - GPT-5.4 mini costs $0.75/$4.5 and is supposed to be the replacement. - GPT-5.4 nano costs $0.2/$1.25 and, while it ranks better in benchmarks than GPT-5 mini, it's not even close when you test it in real scenarios. So you're left being forced to go to GPT 5.4 mini if you use 5 mini today. The same thing is happening here…
Each model release gives an opportunity to reduce the number of old models still on offer, and charge a higher, less-subsidized tier. The trick is to charge a subsidized price that is less than an M3 Ultra, so they continue paying you rent, instead of a one-time fixed cost. So far open models can't compete with Opus 4.5 but as soon as it can, people will be looking at buying devices that can run that model locally. W…
Re: Previewing GPT‑5.6 Sol: a next-generation model
#485If you used GPT-5.5 over the last 24 hours or so, you may have already had access to 5.6. I've been running some tests on a harness we're building, and suddenly saw a jump in a few points yesterday. I reran the vanilla codex benchmark and saw an ~88% score on Terminal Bench 2.1 from GPT-5.5 on vanilla Codex. The biggest indicator, beyond the score, was that 3 tests which frequently hit "safety" blockers with 5.5 star…
[flagged]
Contrary to your predisposition, we're actually quite peeved that we might be seeing results from 5.6 instead of 5.5, as it's muddying our own internal data.
We've run the tasks on this benchmark hundreds of times for our own internal harness. It got magically better yesterday. Last week we were seeing worse performance (sub-80%).
I agree that benchmarks don't mean much for real world use, and I'm a bit disappointed at the lack of variety in the published benchmarks so far.
With that said, 88.8% is higher than Mythos, and the highest I've seen from vanilla Codex. If 5.6 is any better than 5.5, you'd think they would avoid publishing just one coding-related benchmark with a score that equals their previous model.
> I'm not sure why a higher scores on a few tests [..]
It's not just higher scores, the API is no longer flagging tests for cybersecurity warnings that it's been flagging for weeks.
Re: Previewing GPT‑5.6 Sol: a next-generation model
#486If you used GPT-5.5 over the last 24 hours or so, you may have already had access to 5.6. I've been running some tests on a harness we're building, and suddenly saw a jump in a few points yesterday. I reran the vanilla codex benchmark and saw an ~88% score on Terminal Bench 2.1 from GPT-5.5 on vanilla Codex. The biggest indicator, beyond the score, was that 3 tests which frequently hit "safety" blockers with 5.5 star…
these things can just change with infrastructure changes rather than be some mysterious A/B testing.
With that said, I doubt OpenAI would choose to publish a singular coding benchmark for a new model that exactly matches their previous model (88.8%).
Re: Previewing GPT‑5.6 Sol: a next-generation model
#487Re: Previewing GPT‑5.6 Sol: a next-generation model
#488Earlier quoted context omitted.
No offense but have you considered the strong possibility that you’re just not good at what you do? I am occassionally pleased but mostly annoyed or disappointed… but never getting anything close to chills. That sounds downright weird.
No offense but have you considered the strong possibility that you're just holding it wrong? You're entitled to your opinion, but OP is hardly the first person to say something like this and is surrounded by tons of folks saying the exact same thing. Just because it sounds weird to you, doesn't mean it's not true.
Re: Previewing GPT‑5.6 Sol: a next-generation model
#489Re: Previewing GPT‑5.6 Sol: a next-generation model
#490Earlier quoted context omitted.
Why are you representing this as such a binary here? For SLM we don’t need the Taalas stuff at all. Just run it locally on your own device if it’s truly a small model. And there’s plenty of larger models that can be run on-premise just fine. I think it’s impressive that a frontier model can achieve 750t/s. That’s all. You can get similar insane token speeds from other open weight models too.
The irony here is, according to you, my take is the binary one. When your response is: well, we can all just run it on our devices - we don't need any other options! You seem to be cool with a very small and gated ecosystem with whatever tech billionaires want you to have access to. I grew up in the era where compute was diverse and open. You may think this is OK, but it's not. The more options we have and the more d…
Once again, my statement is that the Taalas product is not a fair comparison because it runs an old outdated model. If you want to run a similar model at similar speeds (albeit not serially, but in parallel) you don’t need their product.