Live data from Hacker News

GPT-5.6

openai.com

101–110 of 1001 posts

Re: GPT-5.6

#101
"GPT‑5.6 is available starting today across ChatGPT, Codex, and the OpenAI API. The rollout is starting globally now and will continue gradually toward full availability over the next 24 hours."

Re: GPT-5.6

#103
post #86
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

I personally use opencode so I can swap between models and try different options. I'd say I prefer claude (fable and opus 4.8) so far, but curious to see where gpt 5.6 lands. For personal stuff, I've been pretty happy with chatgpt's $20 plan. I believe it has considerably higher limits than claude's $20 plan, and it's enough for the personal stuff I play with (hermes, and some small coding stuff). Also allows me to k…

The $20 GPT plan with GPT 5.5 lasted me, somehow, exactly one smallish fixup feature

Re: GPT-5.6

#104
> On the Artificial Analysis Coding Agent Index, GPT‑5.6 Sol with max reasoning sets a new state of the art at 80, 2.8 points above Fable 5, while using less than half the output tokens, taking less than half the time, and costing about one-third less.

> That advantage extends across the family: Terra performs just above Fable 5, while Luna outperforms Opus 4.8; each does so in roughly one-third of the time, with about half as many output tokens, and at approximately one-quarter the estimated cost.

Wow. I don't believe it. Every indication and twitter post told me that Fable is much more intelligent than Sol and here we are told that even Terra outperforms Fable?

Not only that, Sol doesn't even come with run time classifiers. So it is even more suspicious.

What's even stranger is that OpenAI is directly referencing a competitor in this direct way.

Re: GPT-5.6

#105

CTRL-F: Fable 15 hits Holy shit. They must be feeling very threatened by Fable if they're spending this much energy talking about it in the release notes for their own model.

yikes - looks like you need to go back to stats school gemini - 13 hits opus - 18 hits So they are more threatened by opus than fable, or are they almost as threatened by gemini as they are by fable?

The second paragraph has four mentions of Fable. I think that makes my case pretty clearly.

Re: GPT-5.6

#106
post #82

Earlier quoted context omitted.

Codex has been good for a long time, more expensive but very focused on efficiency. Working with it feels faster and more to the point than Opus models and I trust it more with long-running jobs. Also regular resets vs being at the whim of Anthropic drama all the time is hella nice.

Codex is cheaper on average no? I think the models are expensive but the token efficiency of the harness itself solves the problem.

Yes that's what I meant, the per token cost is higher but as you say the efficiency levels it out / works slightly in Codex's favour.

Re: GPT-5.6

#107
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

Codex has been good for a long time, more expensive but very focused on efficiency. Working with it feels faster and more to the point than Opus models and I trust it more with long-running jobs. Also regular resets vs being at the whim of Anthropic drama all the time is hella nice.

Anyone know what the deal is with the resets?

Re: GPT-5.6

#110
post #68

The frontier graph on all these benchmark are extremely in favor of 5.6 Sol over Fable, more than the best model comparisons in previous iterations. I'd like to know how cherry-picked this is, and what tests it performed less overwhelmingly in, but I suppose that info is not going to be on this post. If it pans out to be as good as it says, that's great. On the other hand, if this model is not overwhelmingly impressi…

The charts are also extremely difficult to parse. They seem auto-generated. Dataset coloring is atrocious.

Regarding your main point, yes, I agree. My impression (as someone who uses both Codex and Claude Code daily) is that OpenAI does a fair amount of benchmaxxing.

Post reply on HN