Live data from Hacker News

Ask HN: Is Codex with GPT 5.5 Extra High being dumbed down?

news.ycombinator.com

1–10 of 12 posts

Ask HN: Is Codex with GPT 5.5 Extra High being dumbed down?

#1
Hi HN, just want to rant and see if anybody can relate.

The product is not the same as i signed up for few month ago and the same shift i've experienced with Claude Code on Opus 4.6-4.7

The best way to describe the difference is you hire a reliable 'intelligent' tech lead who 'gives a shit' and in some time you eventually get over-confident junior dev that acts as destructive token burner.

No thinking during the process, not even really following instructions.

It's pretty much a binary shift that happens and in my experience can't be cured with prompting.

There is no way to actually tell what's happening under the hood with models but the difference is noticeable from the first reply.

I'm on max, GPT 5.5 extra high, always mindful of context window, use plan mode etc.

Is it the model: dumbing down or quietly route to another model for some reason? Is it the harness? or is it my imagination?

Curious what's your recent experience.

PS: frontier models should add "give-a-shit: max" along with 'thinking: max', 'effort: max' and make them actually work.

Re: Ask HN: Is Codex with GPT 5.5 Extra High being dumbed down?

#3
What I've noticed for the past couple of days is a spike in resource consumption for trivial tasks like examining a commit or listing a directory; it goes around in circles with absurd actions, it doesn't need to diff the files in a directory to search for a single file.

Re: Ask HN: Is Codex with GPT 5.5 Extra High being dumbed down?

#10
I feel like what happens is first they release a giant model. Then they start optimizing the model to increase inference speed and reduce costs.

But then they introduce bugs leading to lots of complaints. Usually the complaints are about models being dumbed down but almost always the labs say they are doing no such thing.

So I am thinking if we believe the labs then they have a very error prone optimization effort going on.

Post reply on HN