Live data from Hacker News

GPT-5.6

openai.com

251–260 of 1001 posts

Re: GPT-5.6

#251
> GPT‑5.6 also introduces more predictable prompt caching, including support for explicit cache breakpoints (opens in a new window) and a 30-minute minimum cache life.

Great to read they are moving away from the 5 minute cache defaults. Hopefully other providers follow soon!

Re: GPT-5.6

#253

Oh man, I love capitalism spoiling us here. I was just enjoying my extra Fable credits, now I'll switch to using 5.6 this weekend. I was planning to ration my Anthropic credits, I guess now I do not have to. And I was half wondering if exactly this would happen: right when Fable usage credits were starting to kick in for people, OAI swoops in and takes the puck. As much the AI craze is crazy, this play by play part i…

top it off with anthropic stressing about the release and resetting usage to 0 for the week just now.

Re: GPT-5.6

#254
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

IMO LMArena is the best benchmark that avoids benchmaxxing https://arena.ai/leaderboard/agent 5.6 isn’t on there yet but Fable leads by a significant margin atm

The results here match up to my real-world experience using these models every day at work and switching between them regularly.

Re: GPT-5.6

#255
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

They blocked Claude from being used in a different harness as well squeezed the usage like crazy. Switched to Codex and haven't cared since.

Between the two the biggest difference by far is ... getting your harness / AGENTS.md / skills / tools set up right.

Re: GPT-5.6

#257

Earlier quoted context omitted.

The naming convention is bizarre and doesn't really mean anything to normies. Trying to pick between "Sol" and "Terra" is like asking the average person if they want the Max or the Ultra chip.

The sun is bigger than earth which is bigger than the moon, it's pretty simple really

Which is cheaper to use? The size euphemism is a really roundabout description vs "Nano" and "Pro" for the layperson.

Re: GPT-5.6

#258
post #234

GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game https://arcprize.org/results/openai-gpt-5-6

Seeing the dramatic differences in scores just going from high to xhigh is just another demonstration of the bitter lesson: Just keep scaling search and learning. We are probably going to need a lot more GPUs.

While I think this is true, remember as we get more efficient we just decide to scale even bigger. So more GPUs, and more efficient.

I agree with the sibling comment, effiency is probably the more important component at this point. We are hitting not just a practical engineering roadblock for scaling with current technology, I think we have definitely hit a financial and logistical roadblock for up scaling with the number of GPUs (on an immediate basis)

Re: GPT-5.6

#259
post #9
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

Use a harness that doesn't lock you into a moat, like OpenCode.

Codex CLI is open source too. I don't think there is a difference.

Re: GPT-5.6

#260
>> approximately 700,000 A100e GPU hours of black-box automated red teaming

Amusing that they use A100e as the reference point to sound impressive. Different ways you could make that conversion, but based on FP4 FLOPs (yes it's disadvantageous to A100, that's the point), that's something like 200hr on a GB300 NVL72 rack.

Not nothing either, but far less astounding sounding than 700k hrs.

Post reply on HN