Live data from Hacker News

GPT-5.6

openai.com

211–220 of 1001 posts

Re: GPT-5.6

#211

Earlier quoted context omitted.

With the exception of Fable which is going away anyway, Codex is better especially after the last couple Opus releases. It’s also no longer slower than Claude. You get much more generous usage from the 20x plan. And you get far better uptime. If benchmarks and early tester impressions are accurate, you also get access to Fable level capability at greater speed and lower cost (included in subscription).

> Fable which is going away anyway $2 says nah. You can't take Fable away in a week where GPT-5.6 and Grok 4.5 launch, if you want to hold on to customers.

The fact that they already extended subscription Fable once would suggest it won’t be solely locked behind API next week, but at the same time it really does look like they are doing everything they can to avoid serving it continuously at scale.

Knowing Anthropic, this unfortunately might end up meaning a quietly quantized Fable on subscription.

Re: GPT-5.6

#212
post #97
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

Codex has arguably been better than Claude Code for months now, but it's flown under the radar because it just didn't capture the same viral marketing effect and OpenAI in general has had more optics / PR issues than Anthropic amongst the online developer crowd. I use the word "better" not in the sense that the underlying GPT models are fundamentally smarter or more intelligent, but rather that as a product Codex is…

Agreed. GPT 5.5 will come up with more straightforward solutions with far fewer tokens than Claude. Also, the usage limits are much more generous for Codex than Claude Code for the same monthly plan.

Re: GPT-5.6

#213
post #72
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

Consensus itself does NOT matter, omp is objectively the best harness for power users yet it has 0 hn posts about it, zero. You're fully free to use and try anything and without caring about what others think is right

The fact that I thought that this was amp misspelled until i someone validate omp and the checked myself indicates it's a subjective assertion at best.

Re: GPT-5.6

#214

Earlier quoted context omitted.

Why would you need a guide for that now? We long had to pick different models (and thinking levels) by task and feel.

The naming convention is bizarre and doesn't really mean anything to normies. Trying to pick between "Sol" and "Terra" is like asking the average person if they want the Max or the Ultra chip.

The sun is bigger than earth which is bigger than the moon, it's pretty simple really

Re: GPT-5.6

#215
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

If you can afford to test it seriously, running both in parallel, it's worth a test to see which you prefer. If you can't, don't bother. You're not likely missing anything since they are close to personal preference with most people I know who have meaningfully tried both preferring Claude

Re: GPT-5.6

#216
post #87

Earlier quoted context omitted.

They do disclose that they scored much lower than Fable on SWEBench Pro, which is a pretty high-quality benchmark. I think it's partially just about what they choose to emphasize...

It's worth noting that OpenAI recently came out saying, "We don't think SWEBench Pro is worth reporting any more" - https://openai.com/index/separating-signal-from-noise-coding...

[flagged]

Re: GPT-5.6

#217
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

> What's the consensus today on codex vs claude code, does it really matter anymore?

Consensus is probably the wrong word for the popular opinions reflected in HN that you might get.

I would recommend that you have 2 of each at all times when it comes to AI so you don't necessarily become overly locked to quirks of one thing. You'll soon realize that things move so fast that you just start internalizing common patterns instead of depending on one specific vendor.

I recommend that you try pi and codex besides claude, to get your own feel for it.

Re: GPT-5.6

#218

GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game https://arcprize.org/results/openai-gpt-5-6

Very interesting. My prediction is that Mythos would outperform Sol.

Also what does this tell about Yann LeCuns whole world model theory? Bro has been going on and on about it. He has made multiple wrong predictions on the trajectory of LLMs.

At some point his claim should be fully falsified no?

Re: GPT-5.6

#219
cursor benchmarks with GPT 5.6 in picture, a good reason to stop using opus.

https://cursor.com/evals

The good news you don't have to send your dollars to China to fund ai dictatorship, in russia, north korea, african countries and south america.

Post reply on HN