Live data from Hacker News

GPT-5.6

openai.com

81–90 of 1001 posts

Re: GPT-5.6

#81
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

Codex UI is way way way better than Claude Code

- codex UI is much more responsive

- i get feedback about the progress easily

- the tool calls and results are very legible, I can click them and see the progress

- no one talks about this but the tool call and response notification are handled much more elegantly in Codex. In Claude Code, it is handled in a clunky way using loops which always causes some delay

- you can steer the conversation midway in Codex

- /side is underrated (/btw is the equivalent and is much worse in Claude Code)

- I have to admit subagents are handled better in Claude Code

Re: GPT-5.6

#82
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

Codex has been good for a long time, more expensive but very focused on efficiency. Working with it feels faster and more to the point than Opus models and I trust it more with long-running jobs. Also regular resets vs being at the whim of Anthropic drama all the time is hella nice.

Codex is cheaper on average no? I think the models are expensive but the token efficiency of the harness itself solves the problem.

Re: GPT-5.6

#83
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

If you can afford it and you have something to justify the expense, I would get both. they're interesting to run side by side, you can hand things off from one to the other. Pretty neat. Unfortunately now I just want to have both :(

Re: GPT-5.6

#84

The meat of the report for SWEs: SWE-Bench Pro Sol: 64.6% Fable: 80% Opus: 69.2% (!!!!) So, it still trails Opus, significantly, and is not a next-gen coding model like Mythos/Fable 5. Disappointing to say the least, but somewhat expected.

SWE-Bench pro is pretty much useless now even though many ppl still look at it. OpenAI published a report yesterday saying so as well. Only look at DeepSWE and FrontierCode right now for coding imo.

Re: GPT-5.6

#85
post #38

Funny to see that they did not include Fable 5 in their GeneBench and LifeSciBench comparisons because "it does not answer advanced biology questions and refuses the majority of questions in this eval". Winner by default!

[deleted]

Re: GPT-5.6

#86
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

I personally use opencode so I can swap between models and try different options. I'd say I prefer claude (fable and opus 4.8) so far, but curious to see where gpt 5.6 lands.

For personal stuff, I've been pretty happy with chatgpt's $20 plan. I believe it has considerably higher limits than claude's $20 plan, and it's enough for the personal stuff I play with (hermes, and some small coding stuff). Also allows me to keep up to date on openai models.

Re: GPT-5.6

#87
post #68

The frontier graph on all these benchmark are extremely in favor of 5.6 Sol over Fable, more than the best model comparisons in previous iterations. I'd like to know how cherry-picked this is, and what tests it performed less overwhelmingly in, but I suppose that info is not going to be on this post. If it pans out to be as good as it says, that's great. On the other hand, if this model is not overwhelmingly impressi…

They do disclose that they scored much lower than Fable on SWEBench Pro, which is a pretty high-quality benchmark. I think it's partially just about what they choose to emphasize...

Re: GPT-5.6

#88

The meat of the report for SWEs: SWE-Bench Pro Sol: 64.6% Fable: 80% Opus: 69.2% (!!!!) So, it still trails Opus, significantly, and is not a next-gen coding model like Mythos/Fable 5. Disappointing to say the least, but somewhat expected.

You've overstated the conclusion. The SWE-bench series has had issues since its inception.

OpenAI no longer recommends SWE-Bench-Pro as a benchmark: https://openai.com/index/separating-signal-from-noise-coding...

Re: GPT-5.6

#89

The meat of the report for SWEs: SWE-Bench Pro Sol: 64.6% Fable: 80% Opus: 69.2% (!!!!) So, it still trails Opus, significantly, and is not a next-gen coding model like Mythos/Fable 5. Disappointing to say the least, but somewhat expected.

[dead]

Re: GPT-5.6

#90
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

There is so much less drama involved with the Codex world. You don't realize how oppressive CC is until you've escaped it. Outages, weird restrictions, degradation, accelerated usage, etc etc etc.

Totally. My experience as well. After some time with codex you're like come on Claude can you just stfu! Haha. I now almost always instruct Claude with specific length requirements when I ask questions. Otherwise, it just blathers and blathers in the most annoying of ways. "Oppressive" is spot on in my opinion
Post reply on HN