Live data from Hacker News

GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

github.com

111–120 of 165 posts

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#111
post #62

I’ve definitely experienced step jumps down in quality on an almost daily basis. I usually used xhigh. The experience of relying on codex’s outstandingly thorough coding earlier in the year has evaporated for me. I’m seeing incredibly stupid implementations intermittently, and have simply switched to Claude until openai takes the issue seriously. As far as i could tell they haven’t taken it seriously for the several…

I have noticed this degradation of 5.5 reliability to what, in my experience, I consider Claude-level of reliability since early June. My journey dealing with this has been transitioning from 5.5 high to 5.5 xhigh to 5.4 high. 5.4 high has been perfectly reliable for me for the last 3 weeks, and I am happy there. Occasionally, I run some tasks on 5.5 xhigh to check if it has gone back to being 100% perfectly reliable…

I'm on the same journey but I bought a 3090 and put qwen 3.6 27b on it. It covers some things with better reliability. Obviously it doesn't have the breadth of a large model. If that's even a selling point for large models for coding?

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#112
post #46

Does this affect the Codex app too, or just the Codex CLI tool?

From some of the numbers I'm seeing in the GitHub issue, the codex desktop app has the same 516 spikes. So most likely it is affected.

If this really is widespread and degrading performance in 40% of the cases, then if OpenAI simultaneously fixes this bug and releases GPT 5.6 within a day or two, then the sudden boost in capability is going to blow people's hair back.

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#113
post #82

Earlier quoted context omitted.

Doesn't look like it: https://marginlab.ai/trackers/codex/

Thanks for sharing this project. Maybe I'm being subjective.

its called hedonic adaptation - you get excited by a new model, but then the excitement disappears, and you confuse that with the model being nerfed

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#114
post #10

Earlier quoted context omitted.

I've switched 3 months ago to Codex because Claude got incredibly stupid. 6 months ago vice versa. It doesn't matter if you use Codex or Claude. Both will fuck with you at some point. Though Codex probably less.

At least OpenAI lets me use my own harness. Having to rely on insane PMs letting Claude Mythos go wild on the codebase has not been going well lately.

Will this problem not arise on other open source harness?

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#115

Deja Vu... This looks just like the Claude Code performance regression back in April. I just quit my Claude subscription when that happened and went to Codex. Now I'm kinda thinking of trying per token for both, using GLM 5.2 on Fireworks for most tasks, shelling out to the big boys only when needed. Not totally confident I'll break even though.

Re per token, I had the same reaction, but given both labs are economically advantaged moving customers to per-token consumption... almost want to avoid this on principle. Even if not intentional, benefitting from a degraded product is not something I want to accept or enable. More now than ever (since original ChatGPT release), the OSS models and open harnesses (eg Pi) are looking mighty attractive.

If pricing is per-token then in theory the vendor can offer you modes that optimize token usage or quality whereas all-you-can eat encourages vendors to satisfy you just enough to keep paying but the trend is towards lower quality responses.

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#116
post #80

Earlier quoted context omitted.

But in that case you have nobody but yourself to blame, and you can stabilize things yourself at any time by refraining from making any changes. You won't be surprised by a provider. Honestly? That's not just valuable—it's essential.

> Honestly? That's not just valuable—it's essential. I'm curious if you wrote this or had a LLM write it. I'm genuinely curious to be clear as I don't see why anyone would bother to go through a LLM to write such a short reply. Have we reached the point where Claudeisms that are this obnoxious have become part of regular speech?

The obnoxious cliche is mine, although I wouldn't call it "regular speech" since I tacked on that dense blob of LLMisms intentionally.

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#117
post #103
post #99

Earlier quoted context omitted.

I’ve always loved em dashes…very sad they’re a hallmark of LLM slop now. That and trios in arguments. Maybe I’m part LLM?

This isn’t just em-dashes—it's the empty phrase that includes both whatever the contrastive construction is called and an “Honestly”. It might have been human written but the density of LLM flags is undeniable.

Honesty has become a big tell, especially if its brutal. I'm not sure if it's getting worse or if I'm becoming more sensitive, but if I'm being brutally honest I can barely stand using LLMs anymore with their tone and their cliches. It's a horrible nightmare fusion of socmed influencer and LinkedIn hustler that sounds freakish even if it were human, and the density of tropes that are accumulating feels ridiculous.

Didn't the foundries take action against those in the past? I don't see "delve" nearly as often anymore. Why are the models spiraling like this now?

What really grinds my gears is the constant need to guess what I'm doing and offer a million random follow-ups. I asked what the weather was like, I don't appreciate the 5 paragraphs of tokens burned on weather-appropriate activity suggestions.

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#119

Earlier quoted context omitted.

You still have to worry about misconfigured local models. Even the professionals get it wrong, which is why local model performance is uneven across providers.

And to add insult to injury, some providers will ride on the good reputation of some local model, selling you a terrible quant instead. With OpenAI, at least my gpt-5.5 is the same as your gpt-5.5. You can't say that about glm for example.

> some providers will ride on the good reputation of some local model, selling you a terrible quant instead.

Quants in popular local inference apps (Ollama, LM Studio, etc) are the worst possible quants (RTN).

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#120

this explains so much why gpt 5.5 has been so bad lately it was really puzzling why it struggled so much where when it first came out it was one shotting stuff totally amazing, i tried the prompt that will tell you if your plan is degraded: codex exec --json --skip-git-repo-check --ephemeral -s read-only --disable memories -m gpt-5.5 -c model_reasoning_effort=high "Do not use external tools. A black bag contains cand…

The correct answer is 29, right? You could draw all the watermelons and all the round pieces before drawing a star piece. So the model never gets it right, but it does when listing the cases exhaustively?
Post reply on HN