GPT-5.6
71–80 of 1001 posts
Re: GPT-5.6
#72Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?
You're fully free to use and try anything and without caring about what others think is right
Re: GPT-5.6
#73Funny to see that they did not include Fable 5 in their GeneBench and LifeSciBench comparisons because "it does not answer advanced biology questions and refuses the majority of questions in this eval". Winner by default!
Re: GPT-5.6
#74Earlier quoted context omitted.
Can't use a claude code subscription in another harness though
You absolutely can; they are not banning anymore. The bigger problem is that subscription versions of the models are way crappier than when the "same" model is hit via API (Bedrock/Vertex) You can also make it not count against extra usage. OpenCode docs show it because Anthropic specifically ambushed them with a PR to remove support so simpletons can't use it easily.
Further, the claim that the subscription "version" of the model is worse sounds like bullshit (and the sort of anecdotal nonsense that you see on sites like this). Do you have anything substantiating this?
Re: GPT-5.6
#75CTRL-F: Fable 15 hits Holy shit. They must be feeling very threatened by Fable if they're spending this much energy talking about it in the release notes for their own model.
gemini - 13 hits
opus - 18 hits
So they are more threatened by opus than fable, or are they almost as threatened by gemini as they are by fable?
Re: GPT-5.6
#76The frontier graph on all these benchmark are extremely in favor of 5.6 Sol over Fable, more than the best model comparisons in previous iterations. I'd like to know how cherry-picked this is, and what tests it performed less overwhelmingly in, but I suppose that info is not going to be on this post. If it pans out to be as good as it says, that's great. On the other hand, if this model is not overwhelmingly impressi…
Re: GPT-5.6
#77Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?
Re: GPT-5.6
#78Re: GPT-5.6
#79Re: GPT-5.6
#80The meat of the report for SWEs: SWE-Bench Pro Sol: 64.6% Fable: 80% Opus: 69.2% (!!!!) So, it still trails Opus, significantly, and is not a next-gen coding model like Mythos/Fable 5. Disappointing to say the least, but somewhat expected.