Live data from Hacker News

GPT-5.6

openai.com

71–80 of 1001 posts

Re: GPT-5.6

#72
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

Consensus itself does NOT matter, omp is objectively the best harness for power users yet it has 0 hn posts about it, zero.

You're fully free to use and try anything and without caring about what others think is right

Re: GPT-5.6

#73
post #38

Funny to see that they did not include Fable 5 in their GeneBench and LifeSciBench comparisons because "it does not answer advanced biology questions and refuses the majority of questions in this eval". Winner by default!

Where’s the lie?

Re: GPT-5.6

#74

Earlier quoted context omitted.

Can't use a claude code subscription in another harness though

You absolutely can; they are not banning anymore. The bigger problem is that subscription versions of the models are way crappier than when the "same" model is hit via API (Bedrock/Vertex) You can also make it not count against extra usage. OpenCode docs show it because Anthropic specifically ambushed them with a PR to remove support so simpletons can't use it easily.

They aren't banning it anymore, they just make it count as "extra usage". e.g. you're paying for every token in addition to your subscription.

Further, the claim that the subscription "version" of the model is worse sounds like bullshit (and the sort of anecdotal nonsense that you see on sites like this). Do you have anything substantiating this?

Re: GPT-5.6

#75

CTRL-F: Fable 15 hits Holy shit. They must be feeling very threatened by Fable if they're spending this much energy talking about it in the release notes for their own model.

yikes - looks like you need to go back to stats school

gemini - 13 hits

opus - 18 hits

So they are more threatened by opus than fable, or are they almost as threatened by gemini as they are by fable?

Re: GPT-5.6

#76
post #68

The frontier graph on all these benchmark are extremely in favor of 5.6 Sol over Fable, more than the best model comparisons in previous iterations. I'd like to know how cherry-picked this is, and what tests it performed less overwhelmingly in, but I suppose that info is not going to be on this post. If it pans out to be as good as it says, that's great. On the other hand, if this model is not overwhelmingly impressi…

The proof is in the pudding and these benchmark stats will only work for so long before people lose interest.

Re: GPT-5.6

#77
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

Codex app is a much different experience than CC CLI. I would try it out for a couple days with the new model suite and see what you prefer after that.

Re: GPT-5.6

#78
I find that 5.5 gives me far fewer refusals than Anthropic models for security and reverse engineering work. I hope the same is true for 5.6.

Re: GPT-5.6

#79
I wish they had kept their previous sensible naming convention instead of this celestial Sol, Terra, and Luna mumbo-jumbo

Re: GPT-5.6

#80

The meat of the report for SWEs: SWE-Bench Pro Sol: 64.6% Fable: 80% Opus: 69.2% (!!!!) So, it still trails Opus, significantly, and is not a next-gen coding model like Mythos/Fable 5. Disappointing to say the least, but somewhat expected.

Makes sense why they released an entire study yesterday discrediting SWE-bench Pro.
Post reply on HN