GPT-5.6
101–110 of 1001 posts
Re: GPT-5.6
#102Re: GPT-5.6
#103Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?
I personally use opencode so I can swap between models and try different options. I'd say I prefer claude (fable and opus 4.8) so far, but curious to see where gpt 5.6 lands. For personal stuff, I've been pretty happy with chatgpt's $20 plan. I believe it has considerably higher limits than claude's $20 plan, and it's enough for the personal stuff I play with (hermes, and some small coding stuff). Also allows me to k…
Re: GPT-5.6
#104> That advantage extends across the family: Terra performs just above Fable 5, while Luna outperforms Opus 4.8; each does so in roughly one-third of the time, with about half as many output tokens, and at approximately one-quarter the estimated cost.
Wow. I don't believe it. Every indication and twitter post told me that Fable is much more intelligent than Sol and here we are told that even Terra outperforms Fable?
Not only that, Sol doesn't even come with run time classifiers. So it is even more suspicious.
What's even stranger is that OpenAI is directly referencing a competitor in this direct way.
Re: GPT-5.6
#105CTRL-F: Fable 15 hits Holy shit. They must be feeling very threatened by Fable if they're spending this much energy talking about it in the release notes for their own model.
yikes - looks like you need to go back to stats school gemini - 13 hits opus - 18 hits So they are more threatened by opus than fable, or are they almost as threatened by gemini as they are by fable?
Re: GPT-5.6
#106Earlier quoted context omitted.
Codex has been good for a long time, more expensive but very focused on efficiency. Working with it feels faster and more to the point than Opus models and I trust it more with long-running jobs. Also regular resets vs being at the whim of Anthropic drama all the time is hella nice.
Codex is cheaper on average no? I think the models are expensive but the token efficiency of the harness itself solves the problem.
Re: GPT-5.6
#107Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?
Codex has been good for a long time, more expensive but very focused on efficiency. Working with it feels faster and more to the point than Opus models and I trust it more with long-running jobs. Also regular resets vs being at the whim of Anthropic drama all the time is hella nice.
Re: GPT-5.6
#108Re: GPT-5.6
#109We have an official pelican on a bicycle from the OpenAI livestream: https://imgshare.cc/mz9xwut3
AGI solved
Re: GPT-5.6
#110The frontier graph on all these benchmark are extremely in favor of 5.6 Sol over Fable, more than the best model comparisons in previous iterations. I'd like to know how cherry-picked this is, and what tests it performed less overwhelmingly in, but I suppose that info is not going to be on this post. If it pans out to be as good as it says, that's great. On the other hand, if this model is not overwhelmingly impressi…
Regarding your main point, yes, I agree. My impression (as someone who uses both Codex and Claude Code daily) is that OpenAI does a fair amount of benchmaxxing.