Live data from Hacker News

GPT-5.6

openai.com

301–310 of 1001 posts

Re: GPT-5.6

#301

Earlier quoted context omitted.

> Fable which is going away anyway $2 says nah. You can't take Fable away in a week where GPT-5.6 and Grok 4.5 launch, if you want to hold on to customers.

The fact that they already extended subscription Fable once would suggest it won’t be solely locked behind API next week, but at the same time it really does look like they are doing everything they can to avoid serving it continuously at scale. Knowing Anthropic, this unfortunately might end up meaning a quietly quantized Fable on subscription.

Can anyone explain this "quietly quantized" model idea to me from a business perspective?

Coca-Cola doesn't "quietly water down" its product to save a few bucks. They know people will take a sip, say "oh that's not what i wanted", and go buy a Pepsi.

If they serve me a quantized Fable, I'm just going to think Fable sucks and go get my tokens elsewhere. What's the point?

Re: GPT-5.6

#302
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

I use Claude for planning, writing CRs, and code review. Codex writes all of the code, no exceptions. Works great, especially when you ask Claude to break up large CRs into roughly 10 minutes of Codex work each.

Same here. I find the design, architecture, system design discussion to be better on Claude, but after Opus 4.6 I switched over to Codex for actual coding and love the results. I use both via the CLI and generally tell Claude to output the result of our decisions as a markdown that will be easy to read and implement by an agentic coding tool. Then I fire up Codex and read said markdown as the input of the session and way to build all the appropriate context needed. I see this as a way to step into letting the agents go run on their own and interact with each other, but I still like to steer so I put these manual steps in the flow. Letting the agents go off on their own and one shot big chunks is not reliable enough yet imo.

Re: GPT-5.6

#303

Earlier quoted context omitted.

> Avoid generic brevity instructions That part is confusing because it's not like they provide an example of how default GPT-5.6 output compares with GPT-5.5 both with default output and prompted for brevity. Whenever I use such prompts, it's usually because I want the model to give me the gist in a few sentences. I'd be stunned if GPT-5.6 was that concise by default. I would think that could "break" a lot of things…

It seems like the way brevity instructions have changed is mis-aligned with how most people would expect to use them or are currently using them. Here's the example they give: > Instead of asking for the shortest possible answer, replace brevity instructions with prioritization: > Lead with the conclusion. Include the evidence needed to support it, any material caveat, and the next action. Omit secondary detail and r…

> Lead with conclusion.

I would presume (perhaps falsely?) that an instruction like this would lead to the model presenting a conclusion not supported by the evidence, and potentially backtracking as it then tries to justify said conclusion.

Yes, if deliberation happens, the model should figure out what it wants to say during that phase; but if you're using auto mode, the model is not going to be doing any deliberating half the time. In those cases, the output blathering is the model's only chance for deliberation. It "thinks as it talks", per se.

Given that, I would advise a different approach: let it blather, but then get it to write you a conclusion at the end that the model can guarantee will obviate the need to read any of the blathering.

I.e. advise the model to add an "executive summary" to the end of any non-trivial-in-length response. With some wording to carefully navigate the model between "the summary is itself too long" vs "the summary acts more like clickbait, leaving out necessary detail such that it requires actually reading the blather."

Not sure exactly what that wording would look like. I imagine something like "write your postscript executive summary as if you were a senior CIA intelligence analyst summarizing ground-level reports into a daily digest for the Joint Chiefs of Staff. Take up as little of their time as possible, but ensure that any detail critical to decision-making is retained." (But that phrasing might only be useful if the model is delivering a certain type of response, and actively counter-productive otherwise. This kind of thing is delicate.)

Re: GPT-5.6

#304

The developer's guide ( https://developers.openai.com/api/docs/guides/latest-model ) has some interesting semantic tips for using the model: > Intent understanding: GPT-5.6 can better infer the user’s underlying goal and intended level of work without you specifying every step. Continue to state important constraints, approval boundaries, and success criteria explicitly. > Original image detail: GPT-5.6 preserves the…

[deleted]

Re: GPT-5.6

#305
post #234

Earlier quoted context omitted.

Seeing the dramatic differences in scores just going from high to xhigh is just another demonstration of the bitter lesson: Just keep scaling search and learning. We are probably going to need a lot more GPUs.

> We are probably going to need a lot more GPUs. Or a breakthrough in algorithms etc. The human brain, heck all bio brains, are proof that you don't need a lot of power or size for intelligence.

20 watts for inference AND training!

Re: GPT-5.6

#306

Earlier quoted context omitted.

I've been using Claude Code, Codex, Gemini (now Antigravity) at the same time for half year now, ever since I dipped my toe into agentic coding. I'd say in general Claude Code and Codex are equally powerful, Gemini is lagging behind. One thing I appreciate with Codex is, OpenAI nowadays sometimes just gives you quota resets you can bank, so when you use up weekly quota before the week ends, you could just reset the q…

I've been using codex app server. Works great. https://learn.chatgpt.com/docs/app-server

Hmm, thanks. Didn't know about this. But looks like a bunch of hassle to set it up?

Re: GPT-5.6

#307

Earlier quoted context omitted.

> Avoid generic brevity instructions That part is confusing because it's not like they provide an example of how default GPT-5.6 output compares with GPT-5.5 both with default output and prompted for brevity. Whenever I use such prompts, it's usually because I want the model to give me the gist in a few sentences. I'd be stunned if GPT-5.6 was that concise by default. I would think that could "break" a lot of things…

It seems like the way brevity instructions have changed is mis-aligned with how most people would expect to use them or are currently using them. Here's the example they give: > Instead of asking for the shortest possible answer, replace brevity instructions with prioritization: > Lead with the conclusion. Include the evidence needed to support it, any material caveat, and the next action. Omit secondary detail and r…

I think instead of "be concise" you could tell it how long the answer should be. I.e. give the answer in one paragraph. Or in 10 lines max.

At least before it would listen to instructions like this.

Re: GPT-5.6

#308

Earlier quoted context omitted.

In my experience, for coding Codex is definitely far ahead of Claude Code, even when using Fable 5 as a model.

you have a very strange experience

People work on different things, fail to mention their field of usage in the comments, and then misunderstand the experiences of others who do the same. Repeat ad nauseam.

Re: GPT-5.6

#309
post #38

Funny to see that they did not include Fable 5 in their GeneBench and LifeSciBench comparisons because "it does not answer advanced biology questions and refuses the majority of questions in this eval". Winner by default!

You shouldn't know too much about biology, stupid human. You might live your life in an unexploitable way.

Re: GPT-5.6

#310

Earlier quoted context omitted.

Pi is so “unbloated” that it’s extra effort to use. You can decide how much work to put into it. I get the trade off. But this is a big jump from CC. I’d recommend some middle ground like opencode.

Even simpler, use Cursor with any frontier model. I see others sweat to add enough context to Claude Code while Cursor has a ton of contextual awareness, uses subagents automatically and is significantly faster with no drop off I have found. I'm not sure why devs are so enamored with living in the CLI, but Cursor has one of those too.

Similar to running arch Linux. Many people do need to. But many people just like tinkering. Tinkering can lead to positive outcomes but it’s usually not “doing work”.
Post reply on HN