Live data from Hacker News

GPT-5.6

openai.com

151–160 of 1001 posts

Re: GPT-5.6

#151
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

It's not clear replies to this thread aren't openAI employees or incentivized influencers, but every benchmark has gpt-5.5 underperforming opus 4.8, sometimes by as much as 10%.

Can they all be wrong/paid-off?

Re: GPT-5.6

#152
I really wish there was just an easy guide on when to use Sol vs Terra vs Luna, and it just moves further into confusing territory when it comes to naming.

The naming convention is especially difficult to decipher depending on what your native language is. Of course a latin language speaker might be able to easily determine oh yeah each one is slightly bigger than the other but I still think it borderlines too confusing.

That aside all the numbers look amazing, and I'll be happy to probably main this alongside grok-4.5 for a while comparing the two on price and efficiency.

I vastly prefer the direction that OpenAI seems to be going with token efficiency and performance compared to Anthropic who seems to be moving towards a world where you just token-max as much as possible ignoring any and all costs.

Re: GPT-5.6

#155
post #84

The meat of the report for SWEs: SWE-Bench Pro Sol: 64.6% Fable: 80% Opus: 69.2% (!!!!) So, it still trails Opus, significantly, and is not a next-gen coding model like Mythos/Fable 5. Disappointing to say the least, but somewhat expected.

SWE-Bench pro is pretty much useless now even though many ppl still look at it. OpenAI published a report yesterday saying so as well. Only look at DeepSWE and FrontierCode right now for coding imo.

Amazing, a company that does poorly in a benchmark says that benchmark is useless...

Re: GPT-5.6

#156
post #29

Earlier quoted context omitted.

Claude Code is a massively bloated agent harness. Try Pi: https://pi.dev/

Pi is so “unbloated” that it’s extra effort to use. You can decide how much work to put into it. I get the trade off. But this is a big jump from CC. I’d recommend some middle ground like opencode.

Even simpler, use Cursor with any frontier model. I see others sweat to add enough context to Claude Code while Cursor has a ton of contextual awareness, uses subagents automatically and is significantly faster with no drop off I have found. I'm not sure why devs are so enamored with living in the CLI, but Cursor has one of those too.

Re: GPT-5.6

#157
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

I sub both codex and claude at 20x. I like opus+fable more than gpt5.5 because it seems gpt tries to finish tasks by leaving any ambiguity unresolved. claude seems better at surfacing open questions.

This is using the same AGENTS.md prompts, which were designed firstly for Claude use, so maybe it's something that could be optimized better if I understood gpt as well?

Re: GPT-5.6

#158
post #68

The frontier graph on all these benchmark are extremely in favor of 5.6 Sol over Fable, more than the best model comparisons in previous iterations. I'd like to know how cherry-picked this is, and what tests it performed less overwhelmingly in, but I suppose that info is not going to be on this post. If it pans out to be as good as it says, that's great. On the other hand, if this model is not overwhelmingly impressi…

[deleted]

Re: GPT-5.6

#159
post #99

We have an official pelican on a bicycle from the OpenAI livestream: https://imgshare.cc/mz9xwut3

holy moly it's in THREE dimensions! AGI solved

So it's failing epically because it generated a tricycle instead of a bicycle?

Re: GPT-5.6

#160
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

> I'm hesitant to leave Claude Code behind for something new.

Codex and Claude Code are not mutually exclusive, you can use both.

Post reply on HN