Live data from Hacker News

GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

github.com

41–50 of 165 posts

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#41
post #16

Oh this seems bad, and is fairly easy to reproduce using codex cli. You give it a puzzle prompt that it has to reason about and solve, occasionally it will seemingly short circuit and think for exactly 516 tokens, and return the wrong result. When it ends up using 6000-8000 thinking tokens it returns the correct result. Maybe some issue with adaptive thinking? Another point for local models I guess, don't have to wor…

[deleted]

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#42

Earlier quoted context omitted.

At least OpenAI lets me use my own harness. Having to rely on insane PMs letting Claude Mythos go wild on the codebase has not been going well lately.

How do you mean OpenAI lets you use your own harness? I'm under the impression that a custom harness requires the OpenAI SDK, which requires api tokens rather than plus/pro accounts.

https://x.com/thsottiaux/status/2058071172361998482

"A little secret. About 5% of our production traffic is on the Pi harness, about another 5% is on OpenCode. Reminder you can use your ChatGPT account in a flourishing set of other tools.

We’ll continue to make Codex awesome, but you have options."

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#43
post #10

Earlier quoted context omitted.

I've switched 3 months ago to Codex because Claude got incredibly stupid. 6 months ago vice versa. It doesn't matter if you use Codex or Claude. Both will fuck with you at some point. Though Codex probably less.

If you use GitHub Copilot you can switch between them in the same session if you want.

Same with a third-party open harness.

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#44

Deja Vu... This looks just like the Claude Code performance regression back in April. I just quit my Claude subscription when that happened and went to Codex. Now I'm kinda thinking of trying per token for both, using GLM 5.2 on Fireworks for most tasks, shelling out to the big boys only when needed. Not totally confident I'll break even though.

Re per token, I had the same reaction, but given both labs are economically advantaged moving customers to per-token consumption... almost want to avoid this on principle. Even if not intentional, benefitting from a degraded product is not something I want to accept or enable.

More now than ever (since original ChatGPT release), the OSS models and open harnesses (eg Pi) are looking mighty attractive.

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#45

Deja Vu... This looks just like the Claude Code performance regression back in April. I just quit my Claude subscription when that happened and went to Codex. Now I'm kinda thinking of trying per token for both, using GLM 5.2 on Fireworks for most tasks, shelling out to the big boys only when needed. Not totally confident I'll break even though.

Fireworks?

Provides access to AI models for a per-token fee. See OpenRouter, they are one of many.

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#47
post #22

Earlier quoted context omitted.

The vibe-assumed claude code performance regression, yep. People should stop expecting consistent performance from non-deterministic systems. There is zero empirical corroboration of performance degredation. There has been a step change... in the amount of whining and complaining coders exhibit lately.

If you bother to look at the issue instead of whining and complaining, you will see the evidence.

This is evidence of a bug, not the purposeful enshittification people are referencing

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#48
post #37
post #23

Earlier quoted context omitted.

I assume they are lying and still think you can use gpt 5.5 non-codex within codex cli. And they outed themselves. A lot of nonsense. And the very poor communication skills just seem like the typical chinese astroturfing you see pretty often now when discussing OAI/Claude.

See, this is part of the confusion. There is no such thing as "GPT-5.5-codex". The last codex-branded model was "GPT-5.3-codex". Starting with "GPT-5.4" the main model handles agentic engineering and they did not release a coding model. Both the web harness and codex app/cli use "GPT-5.5".

haha woops. guess im the chinaman now

Re: GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance

#49
post #48
post #37

Earlier quoted context omitted.

See, this is part of the confusion. There is no such thing as "GPT-5.5-codex". The last codex-branded model was "GPT-5.3-codex". Starting with "GPT-5.4" the main model handles agentic engineering and they did not release a coding model. Both the web harness and codex app/cli use "GPT-5.5".

haha woops. guess im the chinaman now

What do you mean by that? Seems kinda racist.
Post reply on HN