GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance
11–20 of 165 posts
Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance
#12I swear some days ago someone here claimed Openai succeeded cutting down their compute cost by half with a breakthrough optimization. So this is it?
> OpenAI engineers earlier this month told some colleagues they had figured out a way to more than halve the cost of inference, or running existing models, thanks to some newly-discovered optimizations, according to a person with knowledge of those discussions.
Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance
#13Clearly they are batching reasoning inference in a few multiples of 512 tokens as a throughput optimization
Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance
#14I’ve definitely experienced step jumps down in quality on an almost daily basis. I usually used xhigh. The experience of relying on codex’s outstandingly thorough coding earlier in the year has evaporated for me. I’m seeing incredibly stupid implementations intermittently, and have simply switched to Claude until openai takes the issue seriously. As far as i could tell they haven’t taken it seriously for the several…
Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance
#15Personally, I would say very likely, to be honest. I gotta go through this a little more, but I actually use 5.5 codex an obscene amount, and I almost never use it for reasoning anymore. It's not even in the same galaxy as far as actually taking out the thinking and using GPT-5.5 or even Claude and then coming back and giving it the reasoning. Blah blah blah, it's the same model. Well, let me tell you, no, it's not,…
Care to explain what you mean by that?
Not sure if I agree, but I do happen to use a fair bit of web harness as well, just because I find it to be much more effective at web search and a different type of reasoning. So I must agree a little or else I wouldn't do that.
Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance
#16Maybe some issue with adaptive thinking? Another point for local models I guess, don't have to worry about silent server side changes.
Edit: To follow up, it seems to happen quite often. Out of 10 runs of the exact same prompt, 4/10 had this 516 thinking token issue, and every one of these had the wrong solution. So nearly half the time, 5.5 xhigh could be short circuiting and degrading performance. Granted the sample size is small.
Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance
#17Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance
#18Now I'm kinda thinking of trying per token for both, using GLM 5.2 on Fireworks for most tasks, shelling out to the big boys only when needed. Not totally confident I'll break even though.
Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance
#19Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance
#20A rare case "they made the model dumber" where they actually made the model dumber, instead of the usual user psychosis?
I don't find "usual user psychosis" particularly fair or tasteful anyhow. You're not left with much more than subjective judgement and speculation/suspicion when all you have is a magic sink of an API endpoint that ingests your context window then spits back a continuation of it. Even if you have a standardized model test suite, claiming a stealth nerf remains an exercise in mind reading (of the people working there). Model quality can degrade without an explicit intention that way, or a downgrade of the underlying infrastructure, after all.
Being tongue-in-cheek conspiratorial, or even actually entertaining the possibility of a nerf, is no psychosis anyways. Not a fan of this trend of people abusing psychology diagnosis terminology like this. I'm sure there are people who go a step beyond and are overconfident in these judgements, maybe in their case it holds. But then that's a minority, and so what you have then is a hyperboly. Doesn't serve anyone.