Here are 18 pelicans - six each for Luna, Terra and Sol at the six different reasoning effort levels (plus the price to generate each one): https://static.simonwillison.net/static/2026/gpt-5.6-pelican... Or if you want to see some in 3D, OpenAI featured a pelican riding a tricycle, bicycle, pony and another pelican in their livestream this morning: https://www.youtube.com/live/Wq45rvPGNHs?t=1070s
Time to dump this test. Probably not a coincidence every version has the same rolling green hills, gradient blue sky, sun in the corner, etc.
On the one hand: yes, pelicans on bikes are definitely in the training set at this point.
On the other hand: the test is clearly not saturated, given that you can see a clear difference in output at the various reasoning levels / model versions.
Dirac ( https://github.com/dirac-run/dirac , https://dirac.run/ ) now supports gpt-5.6. This thing does now seem to be on the chatGPT/codex accounts yet. UPDATE: it is now available in chatGPT account also, they rolled it out
> We are probably going to need a lot more GPUs. Or a breakthrough in algorithms etc. The human brain, heck all bio brains, are proof that you don't need a lot of power or size for intelligence.
The human brain has 80 billion neurons and a 100 trillion synapses. I think you're underselling the processing power of that warm chunk of meat. The real message of the last 15 years has actually been the opposite: if you throw enough processing power at it, intelligence emerges.
Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?
I recommend trying Codex too. In fact, I recommend running them side-by-side if you have the budget, e.g. have both independently plan the same feature or implement in a different worktree, or have them critique each other's work. I personally find GPT-5.5 to be a better programmer than Opus 4.8, it is extremely thorough, but I don't like the code it generates ("austere"), and find Opus 4.8 to write more "human frien…
I haven't tried Codex yet, but I for my tasks GPT-5.5 may correctly point to a proper direction but its code feels a bit weird. Opus 4.8 is way better in coding, and actually it's the only one who could catch very very sophisticated bug in a large codebase (I tried different models including GPT-5.5 and DeepSeek). Interestingly Gemma 4 under opencode running locally performs not bad at all, it's far yet from DeepSeek level, but it manages to understand tools quite well, and code quality is pretty good. So, for simple coding projects I can say local models already won. It's amazing how smart open models of desktop size have become today. I mean it's quite plausible to manage small codebase today relying on only open tools and local models, you don't need any subscription to produce high quality code, but yes I assume you already experienced and know what you're doing :)
The developer's guide ( https://developers.openai.com/api/docs/guides/latest-model ) has some interesting semantic tips for using the model: > Intent understanding: GPT-5.6 can better infer the user’s underlying goal and intended level of work without you specifying every step. Continue to state important constraints, approval boundaries, and success criteria explicitly. > Original image detail: GPT-5.6 preserves the…
> Avoid generic brevity instructions That part is confusing because it's not like they provide an example of how default GPT-5.6 output compares with GPT-5.5 both with default output and prompted for brevity. Whenever I use such prompts, it's usually because I want the model to give me the gist in a few sentences. I'd be stunned if GPT-5.6 was that concise by default. I would think that could "break" a lot of things…
If you conceptualize this as “there is an appropriate amount of brevity for each situation” then it would be expected for a better model to use different amounts of brevity if it gets better at determining the appropriate amount.
My view is that popular models by default output wildly excessive amounts of prose for nearly every use case, so if this changes in a new model that’s a pure win.
I find that 5.5 gives me far fewer refusals than Anthropic models for security and reverse engineering work. I hope the same is true for 5.6.
Yeah, I pretty much had to switch to using GPT rather than Opus completely for all my security benchmarking and harness development. I was annoyed enough to blog about it: https://swelljoe.com/post/why-i-had-to-switch-to-gpt/
Here are 18 pelicans - six each for Luna, Terra and Sol at the six different reasoning effort levels (plus the price to generate each one): https://static.simonwillison.net/static/2026/gpt-5.6-pelican... Or if you want to see some in 3D, OpenAI featured a pelican riding a tricycle, bicycle, pony and another pelican in their livestream this morning: https://www.youtube.com/live/Wq45rvPGNHs?t=1070s
people are saying this is benchmark is saturated but all of these have occlusion issues, even sol max.
A skilled human artist wouldn't have both legs in front of the bike, or a single straight line representing both leg's crank arms.
Here are 18 pelicans - six each for Luna, Terra and Sol at the six different reasoning effort levels (plus the price to generate each one): https://static.simonwillison.net/static/2026/gpt-5.6-pelican... Or if you want to see some in 3D, OpenAI featured a pelican riding a tricycle, bicycle, pony and another pelican in their livestream this morning: https://www.youtube.com/live/Wq45rvPGNHs?t=1070s
this looks like the same shit from 4 years ago. give it up.
Here are 18 pelicans - six each for Luna, Terra and Sol at the six different reasoning effort levels (plus the price to generate each one): https://static.simonwillison.net/static/2026/gpt-5.6-pelican... Or if you want to see some in 3D, OpenAI featured a pelican riding a tricycle, bicycle, pony and another pelican in their livestream this morning: https://www.youtube.com/live/Wq45rvPGNHs?t=1070s