Live data from Hacker News

GPT-5.6

openai.com

161–170 of 1001 posts

Re: GPT-5.6

#161
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

In my experience, for coding Codex is definitely far ahead of Claude Code, even when using Fable 5 as a model.

you have a very strange experience

Re: GPT-5.6

#162

I really wish there was just an easy guide on when to use Sol vs Terra vs Luna, and it just moves further into confusing territory when it comes to naming. The naming convention is especially difficult to decipher depending on what your native language is. Of course a latin language speaker might be able to easily determine oh yeah each one is slightly bigger than the other but I still think it borderlines too confus…

Why would you need a guide for that now? We long had to pick different models (and thinking levels) by task and feel.

Re: GPT-5.6

#163

The developer's guide ( https://developers.openai.com/api/docs/guides/latest-model ) has some interesting semantic tips for using the model: > Intent understanding: GPT-5.6 can better infer the user’s underlying goal and intended level of work without you specifying every step. Continue to state important constraints, approval boundaries, and success criteria explicitly. > Original image detail: GPT-5.6 preserves the…

> Avoid generic brevity instructions: GPT-5.6 is more sensitive than GPT-5.5 to instructions such as “Be concise,” “Keep it short,” or “Use minimal text.” RIP Caveman skill. Six month good. Now skill dead.

A Yoda skill, is there?

Re: GPT-5.6

#164
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

I use Claude for planning, writing CRs, and code review.

Codex writes all of the code, no exceptions.

Works great, especially when you ask Claude to break up large CRs into roughly 10 minutes of Codex work each.

Re: GPT-5.6

#165
post #66
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

Literally every top model is identical and anyone saying otherwise is engaging in astrology.

anybody saying they're identical clearly doesn't use both...

Re: GPT-5.6

#166
post #66
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

Literally every top model is identical and anyone saying otherwise is engaging in astrology.

The outputs, ui, and overall behavior (tokenization) are not identical.

Re: GPT-5.6

#167
I wish model launches were like proper product releases

it's impossible to _try_ it out on release!

it's not on their codex subscription, or the web/mobile chatgpt interfaces, or aws bedrock, etc. I just cant find a working endpoint with the latest model after they announce

Re: GPT-5.6

#168

Earlier quoted context omitted.

> Avoid generic brevity instructions That part is confusing because it's not like they provide an example of how default GPT-5.6 output compares with GPT-5.5 both with default output and prompted for brevity. Whenever I use such prompts, it's usually because I want the model to give me the gist in a few sentences. I'd be stunned if GPT-5.6 was that concise by default. I would think that could "break" a lot of things…

It seems like the way brevity instructions have changed is mis-aligned with how most people would expect to use them or are currently using them. Here's the example they give: > Instead of asking for the shortest possible answer, replace brevity instructions with prioritization: > Lead with the conclusion. Include the evidence needed to support it, any material caveat, and the next action. Omit secondary detail and r…

Replace 2 word instruction ('be concise') with a 38 word instruction.

Human can no longer be concise when asking for a few sentences instead of 20 paragraphs of BS they don't want to read when all they want is a summary to verify the general direction of the prompt-work before digging into the details.

such progress!

Re: GPT-5.6

#170
post #87
post #68

The frontier graph on all these benchmark are extremely in favor of 5.6 Sol over Fable, more than the best model comparisons in previous iterations. I'd like to know how cherry-picked this is, and what tests it performed less overwhelmingly in, but I suppose that info is not going to be on this post. If it pans out to be as good as it says, that's great. On the other hand, if this model is not overwhelmingly impressi…

They do disclose that they scored much lower than Fable on SWEBench Pro, which is a pretty high-quality benchmark. I think it's partially just about what they choose to emphasize...

SWEBench Pro should be ignored until they fix it or disprove the broken task accusations.
Post reply on HN