Live data from Hacker News

GPT-5.6

openai.com

471–480 of 1001 posts

Re: GPT-5.6

#471

Earlier quoted context omitted.

> Use shorter prompts: In internal evaluations, replacing long, explicit system prompts with minimal prompts improved scores by roughly 10–15%, while reducing total tokens by 41–66% and cost by 33–67%. When has this ever not been the case? I don't think this is a GPT 5.6 specialty!

Information density of the prompt is the most important factor in my experience. And interestingly, LLMs seem particularly bad at writing prompts for other LLMs for this reason (you can guide them to be more dense, just speaking by default). Conciseness is usually a byproduct of information density though.

> LLMs seem particularly bad at writing prompts for other LLMs for this reason

Claude is terrible at this! Probably for the same reason that its writing style in prose is so annoying and full of claudisms.

Re: GPT-5.6

#472
post #379

Here are 18 pelicans - six each for Luna, Terra and Sol at the six different reasoning effort levels (plus the price to generate each one): https://static.simonwillison.net/static/2026/gpt-5.6-pelican... Or if you want to see some in 3D, OpenAI featured a pelican riding a tricycle, bicycle, pony and another pelican in their livestream this morning: https://www.youtube.com/live/Wq45rvPGNHs?t=1070s

They said in the AI community, a pelican riding a bicycle is a good test to measure effectiveness of the model, wondering if they were referring to you, or is it really a standard in the AI community ?

Also would be good to have a tool where users can select models and instantly see each model's generated pelicans. That will make it easy to compare the output of different models.

Re: GPT-5.6

#473
post #379

Here are 18 pelicans - six each for Luna, Terra and Sol at the six different reasoning effort levels (plus the price to generate each one): https://static.simonwillison.net/static/2026/gpt-5.6-pelican... Or if you want to see some in 3D, OpenAI featured a pelican riding a tricycle, bicycle, pony and another pelican in their livestream this morning: https://www.youtube.com/live/Wq45rvPGNHs?t=1070s

They said in the AI community, a pelican riding a bicycle is a good test to measure effectiveness of the model, wondering if they were referring to you, or is it really a standard in the AI community ? Also would be good to have a tool where users can select models and instantly see each model's generated pelicans. That will make it easy to compare the output of different models.

Simon did start the pelicans on bicycles as an SVG, but I think it's more of a fun goofy thing to see how the model performs at. I don't think it has a direct correlation to a model's performance though.

Re: GPT-5.6

#474
post #149
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

Claude Code fan here... Codex is very good. Sometimes better. The killer feature is price. After 6+ months of exclusive Claude Code usage, I was begrudgingly forced to try Codex once Anthropic rejiggered their limits such that I kept maxing out my $200/mo plan in just a few days. These days I pay both $200/mo plans, and it's just about enough to get me through a week's work (small game studio - infinite code to write…

Genuine question/not a critique-are you actually reviewing all that code or just sending it and hoping for the best? I just can't imagine someone is reading/reviewing that much code every day, but maybe I'm wrong?

Re: GPT-5.6

#476
post #234

GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game https://arcprize.org/results/openai-gpt-5-6

Seeing the dramatic differences in scores just going from high to xhigh is just another demonstration of the bitter lesson: Just keep scaling search and learning. We are probably going to need a lot more GPUs.

And dozens of data centers in every state so tokens are dirt cheap.

Re: GPT-5.6

#477
post #234

Earlier quoted context omitted.

Seeing the dramatic differences in scores just going from high to xhigh is just another demonstration of the bitter lesson: Just keep scaling search and learning. We are probably going to need a lot more GPUs.

These aren’t raw base models they are the result of a ton of RLHF and various adjustments. Bitter lesson wildly overstated in this context.

More RLHF is in fact scaling.

Re: GPT-5.6

#478

Earlier quoted context omitted.

> Use shorter prompts: In internal evaluations, replacing long, explicit system prompts with minimal prompts improved scores by roughly 10–15%, while reducing total tokens by 41–66% and cost by 33–67%. When has this ever not been the case? I don't think this is a GPT 5.6 specialty!

Information density of the prompt is the most important factor in my experience. And interestingly, LLMs seem particularly bad at writing prompts for other LLMs for this reason (you can guide them to be more dense, just speaking by default). Conciseness is usually a byproduct of information density though.

Lexical-priming->semantic-space-constraint;specialized-lexis+=sharp distributional-signature;∴ tight concept-cluster; generic-lexis->diffuse-activation, broad candidate-set;Attention-heads key/query-match domain-tokens;"Hamiltonian"->{operator,eigenstate,quantum,energy}->register+domain locked;Net:constrained-decoding,vocab=soft-prior over output-distribution; register-matching;#taskdef=decompress->continue

Re: GPT-5.6

#479
post #149

Earlier quoted context omitted.

Claude Code fan here... Codex is very good. Sometimes better. The killer feature is price. After 6+ months of exclusive Claude Code usage, I was begrudgingly forced to try Codex once Anthropic rejiggered their limits such that I kept maxing out my $200/mo plan in just a few days. These days I pay both $200/mo plans, and it's just about enough to get me through a week's work (small game studio - infinite code to write…

> (small game studio - infinite code to write!) Curious: what multiplier do you think your productivity has increased by, from before AI?

In terms of ability to ship? Easily tenfold. We literally ship 10 times more than before AI. This does not, however, translate into a tenfold increase in actual business success, of course :)

Re: GPT-5.6

#480
post #379

Here are 18 pelicans - six each for Luna, Terra and Sol at the six different reasoning effort levels (plus the price to generate each one): https://static.simonwillison.net/static/2026/gpt-5.6-pelican... Or if you want to see some in 3D, OpenAI featured a pelican riding a tricycle, bicycle, pony and another pelican in their livestream this morning: https://www.youtube.com/live/Wq45rvPGNHs?t=1070s

something is wrong with Terra model series, most pelicans, except Max, looks bad

Seems to match the pareto frontier on Artificial Analysis as well. Terra is nowhere on it.
Post reply on HN