Live data from Hacker News

GPT-5.6

openai.com

141–150 of 1001 posts

Re: GPT-5.6

#141

The developer's guide ( https://developers.openai.com/api/docs/guides/latest-model ) has some interesting semantic tips for using the model: > Intent understanding: GPT-5.6 can better infer the user’s underlying goal and intended level of work without you specifying every step. Continue to state important constraints, approval boundaries, and success criteria explicitly. > Original image detail: GPT-5.6 preserves the…

> Avoid generic brevity instructions That part is confusing because it's not like they provide an example of how default GPT-5.6 output compares with GPT-5.5 both with default output and prompted for brevity. Whenever I use such prompts, it's usually because I want the model to give me the gist in a few sentences. I'd be stunned if GPT-5.6 was that concise by default. I would think that could "break" a lot of things…

It seems like the way brevity instructions have changed is mis-aligned with how most people would expect to use them or are currently using them.

Here's the example they give:

> Instead of asking for the shortest possible answer, replace brevity instructions with prioritization:

> Lead with the conclusion. Include the evidence needed to support it, any material caveat, and the next action. Omit secondary detail and repetition.

> Keep all required facts, decisions, caveats, and next steps. Trim introductions, repetition, generic reassurance, and optional background first.

Generally speaking, when I ask for a short answer, I want a short answer because I'm not really willing to read through a bunch of bullshit to get to a summary. Putting the onus back on me to assume what the model will return and write a longer prompt detailing exactly what information I want completely misses the point of why I'm asking for a short answer in the first place.

Re: GPT-5.6

#142
post #97
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

Codex has arguably been better than Claude Code for months now, but it's flown under the radar because it just didn't capture the same viral marketing effect and OpenAI in general has had more optics / PR issues than Anthropic amongst the online developer crowd. I use the word "better" not in the sense that the underlying GPT models are fundamentally smarter or more intelligent, but rather that as a product Codex is…

I’d argue the opposite. I’ve switched back and forth from one to the other and Opus/Fable has been constantly better than any GPT in my daily work. It’s a bit slower but it does the things right, with as little code as possible, some comments where needed. Codex is faster but you always have to correct it because it got something wrong; it writes tons of code ("let me add a small helper") with obvious comments.

Re: GPT-5.6

#143
post #87

Earlier quoted context omitted.

They do disclose that they scored much lower than Fable on SWEBench Pro, which is a pretty high-quality benchmark. I think it's partially just about what they choose to emphasize...

I totally missed that, because in the charts they showcase for coding, the SWEBench score is not present, they only include it at the end of the post in tables. Hmm. Great catch.

The SWEBench benchmarks are really gamed at this point and should not be trusted period. The solutions are effectively in the training sets and have been for a while.

Re: GPT-5.6

#144

Earlier quoted context omitted.

Anyone know what the deal is with the resets?

They've discovered it's a good marketing strategy. Whenever there's an outage, or a new launch, there's often a reset with it, which helps keep people engaged with OAI / Tibo and reduces churn. They've also introduced banked resets, which are really clever. If you have a $200/month plan and three banked resets, you're not churning because you will overweight giving up those resets (loss aversion theory).

I ran out of resets :( hehe I had 3 and used them all

Re: GPT-5.6

#146
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

There was just a study showing that when presented blindly no one could tell the difference yet users were avid they could

There _is_ a difference in the way Claude and GPT write. Last Friday I felt Opus was becoming dumb because it was writing like GPT.

Re: GPT-5.6

#148

Not available - checked and it's not there.

As usual, even though GPT-5.6 is releasing today, the rollout in ChatGPT and Codex will be gradual over many hours so that we can make sure service remains stable for everyone (same as our previous launches). We usually start with Pro/Enterprise accounts and then work our way down to Plus. We know it's slightly annoying to have to wait a random amount of time, but we do it this way to keep service maximally stable. T…

Understood thanks; will 5.6 fix this issue that makes Pro unusable?

https://github.com/openai/codex/issues/30364

"GPT-5.5 Codex reasoning-token clustering at 516/1034/1552 may be leading to degraded performance on complex tasks"

Re: GPT-5.6

#149
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

Claude Code fan here... Codex is very good. Sometimes better. The killer feature is price.

After 6+ months of exclusive Claude Code usage, I was begrudgingly forced to try Codex once Anthropic rejiggered their limits such that I kept maxing out my $200/mo plan in just a few days. These days I pay both $200/mo plans, and it's just about enough to get me through a week's work (small game studio - infinite code to write!)

Re: GPT-5.6

#150

The way they talk about cyber security fixes makes clear that they are in bed with the government in order to get ahead of Anthropic.

All of them closely collaborate with the government. LLMs are a national security priority and are vetted. Claude AI was used by Palantir's Maven to target the Minab school that led to a triple tap strike killing over 150 schoolchildren.
Post reply on HN