Live data from Hacker News

GPT-5.6

openai.com

171–180 of 1001 posts

Re: GPT-5.6

#171
post #72
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

Consensus itself does NOT matter, omp is objectively the best harness for power users yet it has 0 hn posts about it, zero. You're fully free to use and try anything and without caring about what others think is right

How can this be "objective"? Surely its subjective.

I've tried a fuck load of harnesses but keep coming back to Codex as my harness.

Re: GPT-5.6

#172

I really wish there was just an easy guide on when to use Sol vs Terra vs Luna, and it just moves further into confusing territory when it comes to naming. The naming convention is especially difficult to decipher depending on what your native language is. Of course a latin language speaker might be able to easily determine oh yeah each one is slightly bigger than the other but I still think it borderlines too confus…

Why would you need a guide for that now? We long had to pick different models (and thinking levels) by task and feel.

Previously it was much more obvious which model to reach for depending on your use case because they had the mini and nano naming conventions.

Getting rid of that seems like a step back. Just a personal nit though.

I've seen buzz about this elsewhere as well but to me effort levels seem more like spend limits disguised with another word. I don't think they should even exist.

Re: GPT-5.6

#173
post #87
post #68

The frontier graph on all these benchmark are extremely in favor of 5.6 Sol over Fable, more than the best model comparisons in previous iterations. I'd like to know how cherry-picked this is, and what tests it performed less overwhelmingly in, but I suppose that info is not going to be on this post. If it pans out to be as good as it says, that's great. On the other hand, if this model is not overwhelmingly impressi…

They do disclose that they scored much lower than Fable on SWEBench Pro, which is a pretty high-quality benchmark. I think it's partially just about what they choose to emphasize...

It's worth noting that OpenAI recently came out saying, "We don't think SWEBench Pro is worth reporting any more" - https://openai.com/index/separating-signal-from-noise-coding...

Re: GPT-5.6

#174
post #72

Earlier quoted context omitted.

Consensus itself does NOT matter, omp is objectively the best harness for power users yet it has 0 hn posts about it, zero. You're fully free to use and try anything and without caring about what others think is right

"objectively the best"?

Probably means subjectively according to his own opinion...

Re: GPT-5.6

#175

I really wish there was just an easy guide on when to use Sol vs Terra vs Luna, and it just moves further into confusing territory when it comes to naming. The naming convention is especially difficult to decipher depending on what your native language is. Of course a latin language speaker might be able to easily determine oh yeah each one is slightly bigger than the other but I still think it borderlines too confus…

You don’t know what sol means? You don’t understand the difference in sizes between Terra and sol? I’m genuinely asking.

Re: GPT-5.6

#177

I really wish there was just an easy guide on when to use Sol vs Terra vs Luna, and it just moves further into confusing territory when it comes to naming. The naming convention is especially difficult to decipher depending on what your native language is. Of course a latin language speaker might be able to easily determine oh yeah each one is slightly bigger than the other but I still think it borderlines too confus…

Why would you need a guide for that now? We long had to pick different models (and thinking levels) by task and feel.

The naming convention is bizarre and doesn't really mean anything to normies. Trying to pick between "Sol" and "Terra" is like asking the average person if they want the Max or the Ultra chip.

Re: GPT-5.6

#179

Not available - checked and it's not there.

As usual, even though GPT-5.6 is releasing today, the rollout in ChatGPT and Codex will be gradual over many hours so that we can make sure service remains stable for everyone (same as our previous launches). We usually start with Pro/Enterprise accounts and then work our way down to Plus. We know it's slightly annoying to have to wait a random amount of time, but we do it this way to keep service maximally stable. T…

Is this bug fixed with 5.6? If not, it probably doesn’t matter which version Codex users are getting because the overall result is dramatically worse than stated by Open AI advertising: https://github.com/openai/codex/issues/30364

Re: GPT-5.6

#180
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

I've been using Claude Code, Codex, Gemini (now Antigravity) at the same time for half year now, ever since I dipped my toe into agentic coding. I'd say in general Claude Code and Codex are equally powerful, Gemini is lagging behind. One thing I appreciate with Codex is, OpenAI nowadays sometimes just gives you quota resets you can bank, so when you use up weekly quota before the week ends, you could just reset the q…

Codex is supported well on iPhone/iPad, it’s inside the ChatGPT app.

It’s amazing how much work you can get done on your phone now, especially if you already have a design mapped out in your head.

Post reply on HN