Live data from Hacker News

GPT-5.6

openai.com

131–140 of 1001 posts

Re: GPT-5.6

#131
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

It never really mattered (except when codex was very new). If anything, codex's remote session integration is better, so outside of some "ultracode" orchestration bells/whistles where Claude Code is ahead, I think Codex is a better tool.

Agree, I think there was just a blind study that showed no one could tell the difference even though the users were avid they could

Re: GPT-5.6

#132
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

In my projects, Claude writes and Codex reviews, and I've had a lot of code I've been very happy with out of that, although as of today, Grok _also_ reviews, and finds interesting new stuff.

Re: GPT-5.6

#134
post #16

Most importantly, the cost: > GPT‑5.6 is priced per 1M tokens across three model sizes: Sol is $5 input / $30 output; Terra is $2.50 input / $15 output; and Luna is $1 input / $6 output. Just as expensive as Fable 5. But of course, another slot machine upgrade but the costs will keep going up and the open weight models from china will continue to race everyone else to $0. Looking forward to the next version of GLM, Q…

Also watching deepseek closely. Seems like US frontier labs only know how to throw money at things as opposed to actually make smart improvements to the algorithms.

To be fair, DeepSeek doubled prices during the peak Chinese workday. (Which admittedly doesn't affect me much.)

Re: GPT-5.6

#136

Dirac ( https://github.com/dirac-run/dirac , https://dirac.run/ ) now supports gpt-5.6. This thing does now seem to be on the chatGPT/codex accounts yet. UPDATE: it is now available in chatGPT account also, they rolled it out

Will be there soon according to the last commits in the codex repo: https://github.com/openai/codex/pull/31684/changes

Also, confirmed it works for me by using --model gpt-5.6-sol

Re: GPT-5.6

#137
post #72
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

Consensus itself does NOT matter, omp is objectively the best harness for power users yet it has 0 hn posts about it, zero. You're fully free to use and try anything and without caring about what others think is right

> omp is objectively the best harness for power users

Care to detail this?

Re: GPT-5.6

#138

CTRL-F: Fable 15 hits Holy shit. They must be feeling very threatened by Fable if they're spending this much energy talking about it in the release notes for their own model.

Apparently it significant outperforms fable on both an intelligence and cost index. I don’t believe it at all and I don’t think anyone else does either.

I believe that it outperformed it on benchmarks.

Re: GPT-5.6

#139
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

There was just a study showing that when presented blindly no one could tell the difference yet users were avid they could

Re: GPT-5.6

#140
post #87
post #68

The frontier graph on all these benchmark are extremely in favor of 5.6 Sol over Fable, more than the best model comparisons in previous iterations. I'd like to know how cherry-picked this is, and what tests it performed less overwhelmingly in, but I suppose that info is not going to be on this post. If it pans out to be as good as it says, that's great. On the other hand, if this model is not overwhelmingly impressi…

They do disclose that they scored much lower than Fable on SWEBench Pro, which is a pretty high-quality benchmark. I think it's partially just about what they choose to emphasize...

> SWEBench Pro, which is a pretty high-quality benchmark

No, doesn't seem like it

https://openai.com/index/separating-signal-from-noise-coding...

Post reply on HN