Live data from Hacker News

GPT-5.6

openai.com

351–360 of 1001 posts

Re: GPT-5.6

#351
post #175

Earlier quoted context omitted.

You don’t know what sol means? You don’t understand the difference in sizes between Terra and sol? I’m genuinely asking.

That isn't what "genuinely asking" looks like, you're criticizing using "questions" as cover. It isn't subtle, nor is it constructive. I agree with them, Sol, Terra, and Luna are confusing names. They mean the same thing as GPT-5.6-Max, GPT-5.6-Plus, and GPT-5.6-Fast but require base knowledge for an analogy. It feels like it was adding by the marketing department.

>They mean the same thing as GPT-5.6-Max, GPT-5.6-Plus, and GPT-5.6-Fast but require base knowledge for an analogy.

But do they though? When do you use GPT-5.6-Max-Low vs. GPT-5.6-Plus High? Or GPT-5.6-Fast-Xhigh? What's the Pareto optimal choice (outcome and price)? According to the benches it seems to bop around and the even if the benches are accurate the best choice isn't always consistent.

Re: GPT-5.6

#352
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

The codex software is garbage compared to Claude, but open source is the future, so you should at least switch.

Re: GPT-5.6

#353
I flip back and forth between whoever currently has the more powerful frontier model that isn't cost prohibitive - subscriptions only, API pricing a non-starter. Today that's Fable 5 which has been excellent, as soon as it's Sol I'll switch to that. The OAI/Anthropic harness behavior has mostly stabilized for me with consistent AGENTS.md that I sync with CLAUDE.md - I like pi (pi.dev) and have tried to build it up to get performance comparable to the two "first-party" harnesses, I'm just not there yet.

One major sticking criteria for not going with OpenCode / pi for all of my coding is I want access to the tier-1 frontier model of the day without API pricing - e.g. afaik I can't use Fable 5 via pi harness even though I have a subscription, so for this week I'm on Claude Code. It's not the need to Fable 5 for everything, but even if I just want the marginal intelligence benefit to stress test an architecture decision, it's a safety blanket to know there isn't a ~smarter~ model I could have used. And for my use cases, the doggedness and capability of these frontier models has been insanely effective.

My feeling is we're still in the Uber era subsidy period - the moment the subscriptions either try to lock me in longer than a month or stop OAI/Anthropic stop delivering frontier models in the subscriptions, I'm out - switching fully over to pi.dev or another OS harness and routing my token spend via OpenRouter or offloading to Qwen locally. Then I'll have to put an accurate dollar amount on frontier intelligence.

Re: GPT-5.6

#354

Earlier quoted context omitted.

Wait, what do you mean? 700k A100e hours are equal to 200 hours of a GB300 NVL72 rack? One GB300 NVL72, 72-GPU rack has equal processing power to 3500 A100e GPUs?

maybe? ai says about *8.3 days* of continuous runtime on a single GB300 NVL72 rack about a sprint's level of effort.

a very expensive sprint

Re: GPT-5.6

#355

Just used terra ultra for exactly one prompt in codex and it ate through my full 5h window in about 10mns (20$ plan). The results look pretty good though. Luckily I have had my chatGPT subscription for a while and have a bunch of resets available (nice compared to anthropic). Assuming I take the 5x plan it would give me about an hour of active sessions with terra ultra (maybe ultra is not good value regarding tokens?…

> maybe ultra is not good value regarding tokens?

Well, yes, as explicitly stated on https://openai.com/index/gpt-5-6/: "ultra goes further by coordinating four agents in parallel by default, trading higher token use for stronger results and faster time-to-result on demanding tasks."

Re: GPT-5.6

#356
post #303

Earlier quoted context omitted.

It seems like the way brevity instructions have changed is mis-aligned with how most people would expect to use them or are currently using them. Here's the example they give: > Instead of asking for the shortest possible answer, replace brevity instructions with prioritization: > Lead with the conclusion. Include the evidence needed to support it, any material caveat, and the next action. Omit secondary detail and r…

> Lead with conclusion. I would presume (perhaps falsely?) that an instruction like this would lead to the model presenting a conclusion not supported by the evidence, and potentially backtracking as it then tries to justify said conclusion. Yes, if deliberation happens, the model should figure out what it wants to say during that phase; but if you're using auto mode, the model is not going to be doing any deliberati…

I don't expect that would be the case. This is what's called BLUF or Bottom Line Up Front: https://en.wikipedia.org/wiki/BLUF_(communication)

The model will still have read the entirety of the document before composing its response. And I believe that even in auto mode, there are thinking tokens behind the scenes.

Re: GPT-5.6

#357

Oh man, I love capitalism spoiling us here. I was just enjoying my extra Fable credits, now I'll switch to using 5.6 this weekend. I was planning to ration my Anthropic credits, I guess now I do not have to. And I was half wondering if exactly this would happen: right when Fable usage credits were starting to kick in for people, OAI swoops in and takes the puck. As much the AI craze is crazy, this play by play part i…

Anthropic just reset all limits, including Fable. Capitalism is spoiling us.

Make hay while the sun is out.

Re: GPT-5.6

#358
One of my best use cases for the short duration I have fable is to use it to create the plan and acceptance test files then use GPT 5.5 Pro to do an adversarial review on the plan then feed that feedback into fable to fix the plan.

Re: GPT-5.6

#359
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

It honestly baffles me how people can ask a question like this and get such a wide spectrum of answers in response. It's all so much based on vibes and anecdotal evidence. I've not really noticed much of a difference in capability since Opus 4.6 and I've used a ton of different models. They all work pretty damn well for me.

Re: GPT-5.6

#360

I flip back and forth between whoever currently has the more powerful frontier model that isn't cost prohibitive - subscriptions only, API pricing a non-starter. Today that's Fable 5 which has been excellent, as soon as it's Sol I'll switch to that. The OAI/Anthropic harness behavior has mostly stabilized for me with consistent AGENTS.md that I sync with CLAUDE.md - I like pi (pi.dev) and have tried to build it up to…

I'm working on a multi-harness IDE that supports custom agent workflows and skills that are shared between any harnesses it wraps over. I think it might prove handy for a workflow like yours.
Post reply on HN