Live data from Hacker News

GPT-5.6

openai.com

321–330 of 1001 posts

Re: GPT-5.6

#321
Very interesting: I wonder if the RL approach is diverging between Anthropic and OAI?

I noticed that Fable uses shell tools almost exclusively (even to search and edit files), compared to previous Anthropic models.

Having run some experiments with 5.6, I notice that it uses built-in file systems and provider native tools much more (not shell tools), compared to previous OAI models.

Re: GPT-5.6

#323

The developer's guide ( https://developers.openai.com/api/docs/guides/latest-model ) has some interesting semantic tips for using the model: > Intent understanding: GPT-5.6 can better infer the user’s underlying goal and intended level of work without you specifying every step. Continue to state important constraints, approval boundaries, and success criteria explicitly. > Original image detail: GPT-5.6 preserves the…

> Use shorter prompts: In internal evaluations, replacing long, explicit system prompts with minimal prompts improved scores by roughly 10–15%, while reducing total tokens by 41–66% and cost by 33–67%. When has this ever not been the case? I don't think this is a GPT 5.6 specialty!

For Gemini 2.5 and ~GPT5.0-5.1, longer prompts with lots of explicit instructions and examples produced better conformance. Seems like heavily second guessing the models started to get counter productive around the end of last year.

Re: GPT-5.6

#324
post #234

GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game https://arcprize.org/results/openai-gpt-5-6

Seeing the dramatic differences in scores just going from high to xhigh is just another demonstration of the bitter lesson: Just keep scaling search and learning. We are probably going to need a lot more GPUs.

This isn’t really how it works anymore. Agents rely heavily on tool use and the agentic harness to perform tasks. Pre-training is no longer very effective.

Re: GPT-5.6

#325
post #234

Earlier quoted context omitted.

Seeing the dramatic differences in scores just going from high to xhigh is just another demonstration of the bitter lesson: Just keep scaling search and learning. We are probably going to need a lot more GPUs.

> We are probably going to need a lot more GPUs. Or a breakthrough in algorithms etc. The human brain, heck all bio brains, are proof that you don't need a lot of power or size for intelligence.

For intelligence, I expect the next breakthrough to be colocation of memory and compute in the same chip. And we'll need much more of this memory, probably a few petabytes.

Re: GPT-5.6

#326
For writing GPT which i was subscribed to Fall 2024 to March 2026 (laid off) is superior to Gemini. Been using Gemini since March mostly and they offered a $10 a month plan so i took it. Though today realizing GPT is superior to help me write I am back to being a paying customer. Im in full swing mode to get back into the job market (get the heck away from UI/UX which is now a stupid career in terms of number of jobs out there and in the future there will continue to be less) pivoting into product management (can vibe code anything now) and or customer relations. Hopefully GPT helps me with this pivot and Im again gainfully employed!

Re: GPT-5.6

#328

Earlier quoted context omitted.

I've been using codex app server. Works great. https://learn.chatgpt.com/docs/app-server

Hmm, thanks. Didn't know about this. But looks like a bunch of hassle to set it up?

Finding the right docs/flags took longer than anything else. 15 mins from zero to productive on my phone.

Re: GPT-5.6

#329
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

I prefer codex for most tasks, but stil use Claude if i need to make something "nice but generic", i.e. a html artefact or touch up of front end code.

Re: GPT-5.6

#330
post #284

Earlier quoted context omitted.

They reset all usage half an hour ago. It's back to 0% per week and session. No specifically Fable related.

Hahaha seeing this play out in real time is absolutely incredible.

Im here for it, good on Anthropomorphic to feel some heat again after all that drug dealer Fable business.
Post reply on HN