Live data from Hacker News

GPT-5.6

openai.com

451–460 of 1001 posts

Re: GPT-5.6

#451
post #379

Here are 18 pelicans - six each for Luna, Terra and Sol at the six different reasoning effort levels (plus the price to generate each one): https://static.simonwillison.net/static/2026/gpt-5.6-pelican... Or if you want to see some in 3D, OpenAI featured a pelican riding a tricycle, bicycle, pony and another pelican in their livestream this morning: https://www.youtube.com/live/Wq45rvPGNHs?t=1070s

I think the 'pelican test' is becoming useless. It's been around long enough that now I'm sure good examples are in the training data, and hell they might even do some hand tuning to make it do a decent job since they know people will ask about it.

But either way, with no real way to visualize the result of the text it starts with - it will always be stabbing in the dark. It can't understand conceptually what any of it should look like and then refine the SVG to improve it gradually. It just throws darts at a wall and hopes it comes out alright.

Re: GPT-5.6

#452
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

Set yourself up to be able to try / switch between models easily. I was a claude only user and just have my user level AGENTS.md for codex and others simply point at my user CLAUDE.md. Have a script that syncs my skills (just directories) between all models. Also, if you want to use /simplify or similar from claude in another model, you can ask claude for the prompt and put that in a skill for the other models.

Re: GPT-5.6

#453
post #199

I'm disappointed these models continue to be closed source and so expensive. Open weight models being 10x or more cheaper is just so much more of an unlock than incremental gains for me.

because they're stealing from the frontier models. they're gaming the benchmarks. look how bad glm 5.2 is on cursors evals. gmhit garbage , but it gets glazed as God tier.

Re: GPT-5.6

#454
post #72

Earlier quoted context omitted.

Consensus itself does NOT matter, omp is objectively the best harness for power users yet it has 0 hn posts about it, zero. You're fully free to use and try anything and without caring about what others think is right

"objectively the best"?

To quote a friend:

> Well it's objective _to me_

Re: GPT-5.6

#455

Earlier quoted context omitted.

Same here - gave 5.5 a web design to implement and it sucked. Gave the same to Fable and it still sucked.

did you use Claude design, their tool meant for Web design? because if not then you're the problem .

I did. It was still Claude that’s the problem

Re: GPT-5.6

#456
post #406
post #379

Here are 18 pelicans - six each for Luna, Terra and Sol at the six different reasoning effort levels (plus the price to generate each one): https://static.simonwillison.net/static/2026/gpt-5.6-pelican... Or if you want to see some in 3D, OpenAI featured a pelican riding a tricycle, bicycle, pony and another pelican in their livestream this morning: https://www.youtube.com/live/Wq45rvPGNHs?t=1070s

Ok, I'll never use max effort again on OAI models..

Is that... an x-rated, censored pelican?

Re: GPT-5.6

#457
post #273
post #142

Earlier quoted context omitted.

I’d argue the opposite. I’ve switched back and forth from one to the other and Opus/Fable has been constantly better than any GPT in my daily work. It’s a bit slower but it does the things right, with as little code as possible, some comments where needed. Codex is faster but you always have to correct it because it got something wrong; it writes tons of code ("let me add a small helper") with obvious comments.

I really love the Opus/Fable models but I'm honestly sick to death of the buggy product. The CLI always has some weird issue. Right now it doesn't even output messages before tool calls, it just swallows them and they disappear. I don't like OpenAI as a company, but they appear to have QA, and that is probably enough to get me to switch.

Glad I’m not the only one noticing this. It’s maddening.

Re: GPT-5.6

#458
post #379

Here are 18 pelicans - six each for Luna, Terra and Sol at the six different reasoning effort levels (plus the price to generate each one): https://static.simonwillison.net/static/2026/gpt-5.6-pelican... Or if you want to see some in 3D, OpenAI featured a pelican riding a tricycle, bicycle, pony and another pelican in their livestream this morning: https://www.youtube.com/live/Wq45rvPGNHs?t=1070s

I think the 'pelican test' is becoming useless. It's been around long enough that now I'm sure good examples are in the training data, and hell they might even do some hand tuning to make it do a decent job since they know people will ask about it. But either way, with no real way to visualize the result of the text it starts with - it will always be stabbing in the dark. It can't understand conceptually what any of…

I think it's still useful in a "hello world" sort of way. It means you actually tried the new model.

Re: GPT-5.6

#459

Earlier quoted context omitted.

Surely "how to draw a SVG pelican on a bike" has made it into the training data by now ...

If that was the case in a non-trivial way you'd see mode collapse, but you don't, they come out differently.

It's because all the frequent comment that this pelican is in the trainings set now also got into the trainings set and models adapt. /joking (I hope)

Re: GPT-5.6

#460
post #175

Earlier quoted context omitted.

You don’t know what sol means? You don’t understand the difference in sizes between Terra and sol? I’m genuinely asking.

That isn't what "genuinely asking" looks like, you're criticizing using "questions" as cover. It isn't subtle, nor is it constructive. I agree with them, Sol, Terra, and Luna are confusing names. They mean the same thing as GPT-5.6-Max, GPT-5.6-Plus, and GPT-5.6-Fast but require base knowledge for an analogy. It feels like it was adding by the marketing department.

Surely it's size based: Sun (Sol) > Earth (Terra) > Moon (Luna)

Similar to Anthropic's size/length based naming: Opus > Sonnet > Haiku

These names seem easy to understand to me, and much clearer than suffixes like -max and -plus.

Post reply on HN