Live data from Hacker News

GPT-5.6

openai.com

811–820 of 1001 posts

Re: GPT-5.6

#814

I flip back and forth between whoever currently has the more powerful frontier model that isn't cost prohibitive - subscriptions only, API pricing a non-starter. Today that's Fable 5 which has been excellent, as soon as it's Sol I'll switch to that. The OAI/Anthropic harness behavior has mostly stabilized for me with consistent AGENTS.md that I sync with CLAUDE.md - I like pi (pi.dev) and have tried to build it up to…

> My feeling is we're still in the Uber era subsidy period I often wonder whether this doesn't continue indefinitely. Uber was able to do this because it was just them and Lyft playing second fiddle, with a huge barrier to entry once the network effects had kicked in. It just seems like the model space has way too many competitors, + OSS/Local options for them to ever be able to jack up their prices. At least once th…

Cars are pretty mature while AI is just getting started. Expect 100x price drop for the same quality.

Re: GPT-5.6

#815

I flip back and forth between whoever currently has the more powerful frontier model that isn't cost prohibitive - subscriptions only, API pricing a non-starter. Today that's Fable 5 which has been excellent, as soon as it's Sol I'll switch to that. The OAI/Anthropic harness behavior has mostly stabilized for me with consistent AGENTS.md that I sync with CLAUDE.md - I like pi (pi.dev) and have tried to build it up to…

Zed IDE is allowing me to use Fable with Max x5 or any OpenAI model.

Re: GPT-5.6

#816
From my first tests today, it is a workhouse. It can scan my whole code base, optimize every part, with a greater level of autonomy than other tools. This is insane, we are living at the best time.

Re: GPT-5.6

#817

From my first tests today, it is a workhouse. It can scan my whole code base, optimize every part, with a greater level of autonomy than other tools. This is insane, we are living at the best time.

The literal worst time. I prefer meritocracies.

Re: GPT-5.6

#819
post #667

I use both Claude and Codex, but mostly Claude for planning and coding, and Codex to review Claude’s work. I follow a sort of waterfall workflow which is verbose but fully transparent. Anthropic’s $100 subscription works fine for me, but whatever subscription my company has with OpenAI reaches the 5hr limit ridiculously quickly.

How do you couple them together efficiently? The nice thing about Codex or Claude is that the delegation or multi agent workflow capabilities are just built-in. Do you link one with the other as a skill or mcp or so?

[flagged]

Re: GPT-5.6

#820
post #597

GPT-5.6 is a really good model, and quite cheap. I can finally replace GPT-5.3-Codex for my Tool Calling in n8n. Here's my benchmark results for GPT-5.6: https://aibenchy.com/?q=gpt-5.6 (the high reasoning variants are still running, uploading them soon too) EDIT: The high variants are there too, enjoy the hamsters[0]. [0]: https://aibenchy.com/showcase/?q=gpt-5.6

Given that both Gemini 3.5 Flash (high) and Gemini 3 Flash Preview (medium) beat GPT-5.6 Sol (high) for correctness and score in your benchmarks I don’t trust them at all. The rest of the ranking also doesn’t make sense, like GPT-5.3-Codex (medium) performs better than Claude Opus 4.8 (medium) yeah sure
Post reply on HN