Live data from Hacker News

GPT-5.6

openai.com

541–550 of 1001 posts

Re: GPT-5.6

#541
post #463
post #379

Here are 18 pelicans - six each for Luna, Terra and Sol at the six different reasoning effort levels (plus the price to generate each one): https://static.simonwillison.net/static/2026/gpt-5.6-pelican... Or if you want to see some in 3D, OpenAI featured a pelican riding a tricycle, bicycle, pony and another pelican in their livestream this morning: https://www.youtube.com/live/Wq45rvPGNHs?t=1070s

[flagged]

You can't post like this to HN, regardless of how you feel about someone else's posts.

Doing it repeatedly crosses into harassment, and you've done it more than 3 times now - e.g.

https://news.ycombinator.com/item?id=48504823

https://news.ycombinator.com/item?id=48504654

So please especially stop doing that. If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.

Re: GPT-5.6

#542
post #482

Dirac ( https://github.com/dirac-run/dirac , https://dirac.run/ ) now supports gpt-5.6. This thing does now seem to be on the chatGPT/codex accounts yet. UPDATE: it is now available in chatGPT account also, they rolled it out

Does it support subagents?

yup it does

Re: GPT-5.6

#543
does anyone on chatgpt business plan (not enterprise) not have access to the Sol models in codex? i have 5.6 for terra and luna but not sol

Re: GPT-5.6

#544
post #493
post #439

I love testing the new models by asking them to code a toy RTS game. Here's what Terra did: https://senko.net/vibecode-bench/2026/rts-gpt-5.6-terra.html (one try, in codex app, xhigh effort) Comparing this to other models, I find it similar to GPT-5.5 and a bit behind Sonnet 5. You can see how other models fared here: https://senko.net/vibecode-bench/ (you can also fetch the prompt and the the 5.6 Terra resulting cod…

So the measure of a model is how well they can recreate something they easily have thousands of examples of in their training data. There's probably a better base RTS on github somewhere for free.

The goalpost velocity is approaching light speed…

Re: GPT-5.6

#545
post #463
post #379

Here are 18 pelicans - six each for Luna, Terra and Sol at the six different reasoning effort levels (plus the price to generate each one): https://static.simonwillison.net/static/2026/gpt-5.6-pelican... Or if you want to see some in 3D, OpenAI featured a pelican riding a tricycle, bicycle, pony and another pelican in their livestream this morning: https://www.youtube.com/live/Wq45rvPGNHs?t=1070s

[flagged]

[deleted]

Re: GPT-5.6

#546
post #219

cursor benchmarks with GPT 5.6 in picture, a good reason to stop using opus. https://cursor.com/evals The good news you don't have to send your dollars to China to fund ai dictatorship, in russia, north korea, african countries and south america.

so the answer is use grok ?

I'd say answer , the opus is no longer undisputed. grok + gpt models are very competitive + glm if you are ok to wait 3-4 times longer, unless you have some unique access to GPU

Re: GPT-5.6

#547

GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game https://arcprize.org/results/openai-gpt-5-6

it seems the older models were capped at 10kusd for the runs though?

Re: GPT-5.6

#548
post #142
post #97

Earlier quoted context omitted.

Codex has arguably been better than Claude Code for months now, but it's flown under the radar because it just didn't capture the same viral marketing effect and OpenAI in general has had more optics / PR issues than Anthropic amongst the online developer crowd. I use the word "better" not in the sense that the underlying GPT models are fundamentally smarter or more intelligent, but rather that as a product Codex is…

I’d argue the opposite. I’ve switched back and forth from one to the other and Opus/Fable has been constantly better than any GPT in my daily work. It’s a bit slower but it does the things right, with as little code as possible, some comments where needed. Codex is faster but you always have to correct it because it got something wrong; it writes tons of code ("let me add a small helper") with obvious comments.

I have both as well. I trust the output of Claude to a higher degree than what I get with Codex. I always have claude review codex output. That being said, I find gpt 5.5 more generically useful at a wider breadth of tasks. Straight coding though, it's no contest.

Obligatory YMMV, maybe your prompting style fits gpt better. We forget that this matters a lot

Re: GPT-5.6

#549

The developer's guide ( https://developers.openai.com/api/docs/guides/latest-model ) has some interesting semantic tips for using the model: > Intent understanding: GPT-5.6 can better infer the user’s underlying goal and intended level of work without you specifying every step. Continue to state important constraints, approval boundaries, and success criteria explicitly. > Original image detail: GPT-5.6 preserves the…

Serious question: what is a short prompt?

(For that matter at what point is it "long"? And does the rest of the context matter? Should it be short too?)

Re: GPT-5.6

#550

Earlier quoted context omitted.

That isn't what "genuinely asking" looks like, you're criticizing using "questions" as cover. It isn't subtle, nor is it constructive. I agree with them, Sol, Terra, and Luna are confusing names. They mean the same thing as GPT-5.6-Max, GPT-5.6-Plus, and GPT-5.6-Fast but require base knowledge for an analogy. It feels like it was adding by the marketing department.

Surely it's size based: Sun (Sol) > Earth (Terra) > Moon (Luna) Similar to Anthropic's size/length based naming: Opus > Sonnet > Haiku These names seem easy to understand to me, and much clearer than suffixes like -max and -plus.

I'd agree it is similar to Anthropic's naming scheme, which I'd argue shares the same problems as this. It improves marketability/googlability, but decreases actual comprehension.

You don't actually explain why or how these names are "easy to understand" just state that they simply are. That's great; to me, they aren't obvious or intuitive at all. May have well just start randomly pointing at dictionary words.

Post reply on HN