Live data from Hacker News

GPT-5.6

openai.com

231–240 of 1001 posts

Re: GPT-5.6

#231
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

Not sure about the consensus, but during an entire week I have done every task on my workplace with both Opus 4.8 and GPT 5.5. GPT won hands down. I would even sometimes copy the plans and solutions (using different Git worktrees) from GPT and paste it on Opus and itself would say GPT plans were better. At that point I have migrated. Fable is not enabled in our workspace so I have not tried.

Claude lost my trust around February this year when the plan would say nonsensical things as "delete this method" that was clearly a key method on that part of the codebase.

For personal projects I am using Codex 20$ plan and when that is over I use DeepSeek which is insanely good for the cost.

Re: GPT-5.6

#233
post #142
post #97

Earlier quoted context omitted.

Codex has arguably been better than Claude Code for months now, but it's flown under the radar because it just didn't capture the same viral marketing effect and OpenAI in general has had more optics / PR issues than Anthropic amongst the online developer crowd. I use the word "better" not in the sense that the underlying GPT models are fundamentally smarter or more intelligent, but rather that as a product Codex is…

I’d argue the opposite. I’ve switched back and forth from one to the other and Opus/Fable has been constantly better than any GPT in my daily work. It’s a bit slower but it does the things right, with as little code as possible, some comments where needed. Codex is faster but you always have to correct it because it got something wrong; it writes tons of code ("let me add a small helper") with obvious comments.

Sounds like you are talking past each other. GP is saying the harness of codex is higher quality, which I can believe, even if the models are not as good as Opus/Fable.

Re: GPT-5.6

#234

GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game https://arcprize.org/results/openai-gpt-5-6

Seeing the dramatic differences in scores just going from high to xhigh is just another demonstration of the bitter lesson: Just keep scaling search and learning. We are probably going to need a lot more GPUs.

Re: GPT-5.6

#236

The developer's guide ( https://developers.openai.com/api/docs/guides/latest-model ) has some interesting semantic tips for using the model: > Intent understanding: GPT-5.6 can better infer the user’s underlying goal and intended level of work without you specifying every step. Continue to state important constraints, approval boundaries, and success criteria explicitly. > Original image detail: GPT-5.6 preserves the…

> Use shorter prompts: In internal evaluations, replacing long, explicit system prompts with minimal prompts improved scores by roughly 10–15%, while reducing total tokens by 41–66% and cost by 33–67%.

When has this ever not been the case? I don't think this is a GPT 5.6 specialty!

Re: GPT-5.6

#237
post #230

The developer's guide ( https://developers.openai.com/api/docs/guides/latest-model ) has some interesting semantic tips for using the model: > Intent understanding: GPT-5.6 can better infer the user’s underlying goal and intended level of work without you specifying every step. Continue to state important constraints, approval boundaries, and success criteria explicitly. > Original image detail: GPT-5.6 preserves the…

> Use shorter prompts: In internal evaluations, replacing long, explicit system prompts with minimal prompts improved scores by roughly 10–15%, while reducing total tokens by 41–66% and cost by 33–67%. A shorter prompt results in half as much tokens spend? I find this very hard to believe.

Maybe Codex has the same problem I sometimes have focusing while reading and has to reread the same sentence over and over again.

Re: GPT-5.6

#238
post #66

Earlier quoted context omitted.

Literally every top model is identical and anyone saying otherwise is engaging in astrology.

anybody saying they're identical clearly doesn't use both...

Honestly, I would even push it further. People who would claim that don't use either one.

Re: GPT-5.6

#240
Oh man, I love capitalism spoiling us here. I was just enjoying my extra Fable credits, now I'll switch to using 5.6 this weekend. I was planning to ration my Anthropic credits, I guess now I do not have to. And I was half wondering if exactly this would happen: right when Fable usage credits were starting to kick in for people, OAI swoops in and takes the puck. As much the AI craze is crazy, this play by play part is pretty fun.
Post reply on HN