Live data from Hacker News

GPT-5.6

openai.com

91–100 of 1001 posts

Re: GPT-5.6

#91
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

I use both especially for checking each others work. Pretty happy with results

Re: GPT-5.6

#92
post #80

The meat of the report for SWEs: SWE-Bench Pro Sol: 64.6% Fable: 80% Opus: 69.2% (!!!!) So, it still trails Opus, significantly, and is not a next-gen coding model like Mythos/Fable 5. Disappointing to say the least, but somewhat expected.

Makes sense why they released an entire study yesterday discrediting SWE-bench Pro.

And they'd be right, it's an almost saturated benchmark where even some subpar open source models score very well on. And most models are clustered within a small range so it really doesn't tell you much.

Re: GPT-5.6

#93

Not available - checked and it's not there.

As usual, even though GPT-5.6 is releasing today, the rollout in ChatGPT and Codex will be gradual over many hours so that we can make sure service remains stable for everyone (same as our previous launches). We usually start with Pro/Enterprise accounts and then work our way down to Plus. We know it's slightly annoying to have to wait a random amount of time, but we do it this way to keep service maximally stable.

The timescale is typically hours not minutes, so if you don't see it now, I'd try again later today.

We mention it will be a gradual rollout over the next 24 hours in the Availability section at the bottom of the blog but I admit it's pretty buried.

(I work at OpenAI.)

Re: GPT-5.6

#94

The developer's guide ( https://developers.openai.com/api/docs/guides/latest-model ) has some interesting semantic tips for using the model: > Intent understanding: GPT-5.6 can better infer the user’s underlying goal and intended level of work without you specifying every step. Continue to state important constraints, approval boundaries, and success criteria explicitly. > Original image detail: GPT-5.6 preserves the…

> Avoid generic brevity instructions

That part is confusing because it's not like they provide an example of how default GPT-5.6 output compares with GPT-5.5 both with default output and prompted for brevity. Whenever I use such prompts, it's usually because I want the model to give me the gist in a few sentences. I'd be stunned if GPT-5.6 was that concise by default. I would think that could "break" a lot of things for developers who didn't know to make prompt changes after upgrading to 5.6. What if you were expecting GPT to be as wordy as it usually is? Then suddenly your output is not wordy enough?

Smells like OpenAI trying its best to stave off financial armageddon for another few months. Then again, I'm not sure why they chose to waste so much output computation on verbal diarrhea all this time up to now.

Re: GPT-5.6

#95
post #9

Earlier quoted context omitted.

Use a harness that doesn't lock you into a moat, like OpenCode.

Can't use a claude code subscription in another harness though

You can however for now use wrappers which are not harnesses such as T3Code though. They were going to cut under the Programmatic API, but have at least temporarily walked it back.

Re: GPT-5.6

#96

The developer's guide ( https://developers.openai.com/api/docs/guides/latest-model ) has some interesting semantic tips for using the model: > Intent understanding: GPT-5.6 can better infer the user’s underlying goal and intended level of work without you specifying every step. Continue to state important constraints, approval boundaries, and success criteria explicitly. > Original image detail: GPT-5.6 preserves the…

[deleted]

Re: GPT-5.6

#97
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

Codex has arguably been better than Claude Code for months now, but it's flown under the radar because it just didn't capture the same viral marketing effect and OpenAI in general has had more optics / PR issues than Anthropic amongst the online developer crowd. I use the word "better" not in the sense that the underlying GPT models are fundamentally smarter or more intelligent, but rather that as a product Codex is just simpler, cheaper, and abundantly reliable and low-drama.

Re: GPT-5.6

#98
There is an issue on the page that causes the benchmark tables to get cut off. If you highlight and drag right you can see a few more models like Gemini and Claude Opus. It's also interesting that they introduced explicit caching, which is something that only Anthropic had for a long time.

Re: GPT-5.6

#100

where is it? Still not accessible...

"GPT‑5.6 is available starting today across ChatGPT, Codex, and the OpenAI API. The rollout is starting globally now and will continue gradually toward full availability over the next 24 hours."
Post reply on HN