Live data from Hacker News

GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

github.com

81–90 of 165 posts

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#81
post #79

The good experience I had with GPT-5.5 before made me upgrade to Pro this month. Now I want a refund.

You want a refund because of a problem you weren't even aware of until now? And you don't even really know if your work has been impacted by this problem.

No, the decline in GPT-5.5's performance over the past few weeks is clearly noticeable.

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#82
post #79

Earlier quoted context omitted.

You want a refund because of a problem you weren't even aware of until now? And you don't even really know if your work has been impacted by this problem.

No, the decline in GPT-5.5's performance over the past few weeks is clearly noticeable.

Doesn't look like it: https://marginlab.ai/trackers/codex/

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#83
post #80

Earlier quoted context omitted.

You still have to worry about misconfigured local models. Even the professionals get it wrong, which is why local model performance is uneven across providers.

But in that case you have nobody but yourself to blame, and you can stabilize things yourself at any time by refraining from making any changes. You won't be surprised by a provider. Honestly? That's not just valuable—it's essential.

> Honestly? That's not just valuable—it's essential.

I'm curious if you wrote this or had a LLM write it.

I'm genuinely curious to be clear as I don't see why anyone would bother to go through a LLM to write such a short reply. Have we reached the point where Claudeisms that are this obnoxious have become part of regular speech?

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#84

Maybe its just bad memory but I feel like 5.3 was the best version in terms of token usage and code quality. 5.5 works better but it just eviscerates tokens.

It’s not just you this is also my opinion, 5.3-codex was a fantastic model in terms of balancing output quality and cost.

Cheap and efficient enough I could afford to use it on basically everything unlike 5.5 or Opus, but still pretty good, I preferred it to sonnet

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#85
post #80

Earlier quoted context omitted.

But in that case you have nobody but yourself to blame, and you can stabilize things yourself at any time by refraining from making any changes. You won't be surprised by a provider. Honestly? That's not just valuable—it's essential.

> Honestly? That's not just valuable—it's essential. I'm curious if you wrote this or had a LLM write it. I'm genuinely curious to be clear as I don't see why anyone would bother to go through a LLM to write such a short reply. Have we reached the point where Claudeisms that are this obnoxious have become part of regular speech?

Or they were making a joke[0].

[0]: https://en.wikipedia.org/wiki/Joke

(…just like that)

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#86

Earlier quoted context omitted.

> Honestly? That's not just valuable—it's essential. I'm curious if you wrote this or had a LLM write it. I'm genuinely curious to be clear as I don't see why anyone would bother to go through a LLM to write such a short reply. Have we reached the point where Claudeisms that are this obnoxious have become part of regular speech?

Or they were making a joke[0]. [0]: https://en.wikipedia.org/wiki/Joke (…just like that)

Sure, but it doesn't really fit there as a joke, it looks like it's just meant to be part of what they were trying to say.

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#87
post #70

Even without stats i know it went bad. In the pass two month barely can do any good scientific writing lately, which of course rely on reasoning. It just writing for gods sake. And it show how far we are from AGI.

This is an intermittent issue, you should still be able to get your work done. 5.5 was released two months ago so perhaps you're using 5.5 wrong and some things that worked in 5.4 require tweaking your prompts?

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#88

I’ve definitely experienced step jumps down in quality on an almost daily basis. I usually used xhigh. The experience of relying on codex’s outstandingly thorough coding earlier in the year has evaporated for me. I’m seeing incredibly stupid implementations intermittently, and have simply switched to Claude until openai takes the issue seriously. As far as i could tell they haven’t taken it seriously for the several…

[dead]

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#89
post #80

Earlier quoted context omitted.

But in that case you have nobody but yourself to blame, and you can stabilize things yourself at any time by refraining from making any changes. You won't be surprised by a provider. Honestly? That's not just valuable—it's essential.

> Honestly? That's not just valuable—it's essential. I'm curious if you wrote this or had a LLM write it. I'm genuinely curious to be clear as I don't see why anyone would bother to go through a LLM to write such a short reply. Have we reached the point where Claudeisms that are this obnoxious have become part of regular speech?

I've noticed them trying to creep into my writing. It doesn't help that I was a heavy em-dash user ten years before GPT-3.

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#90

this explains so much why gpt 5.5 has been so bad lately it was really puzzling why it struggled so much where when it first came out it was one shotting stuff totally amazing, i tried the prompt that will tell you if your plan is degraded: codex exec --json --skip-git-repo-check --ephemeral -s read-only --disable memories -m gpt-5.5 -c model_reasoning_effort=high "Do not use external tools. A black bag contains cand…

Verified this locally myself. Thanks for the concrete test. I guess it's time to give Claude another try.
Post reply on HN