> Reload to try again, or go back.
This on iOS, safari
671–680 of 1001 posts
> Reload to try again, or go back.
This on iOS, safari
We Openly hate OpenAI because they’re not very Open but we secretly hope they win against not-open-at-all Anthropic.
I use both Claude and Codex, but mostly Claude for planning and coding, and Codex to review Claude’s work. I follow a sort of waterfall workflow which is verbose but fully transparent. Anthropic’s $100 subscription works fine for me, but whatever subscription my company has with OpenAI reaches the 5hr limit ridiculously quickly.
GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game https://arcprize.org/results/openai-gpt-5-6
Replay of Sol attempting the game: https://arcprize.org/replay/83543d22-8e1e-439a-8809-129ff1d9...
It seems a weird and arbitrary challenge for a language model to be expected to perform. It also seems like there are some harness/visual issues even in the first few steps, where it states that it hasn't moved when it clearly has.
Here are 18 pelicans - six each for Luna, Terra and Sol at the six different reasoning effort levels (plus the price to generate each one): https://static.simonwillison.net/static/2026/gpt-5.6-pelican... Or if you want to see some in 3D, OpenAI featured a pelican riding a tricycle, bicycle, pony and another pelican in their livestream this morning: https://www.youtube.com/live/Wq45rvPGNHs?t=1070s
Time to dump this test. Probably not a coincidence every version has the same rolling green hills, gradient blue sky, sun in the corner, etc.
I really wish there was just an easy guide on when to use Sol vs Terra vs Luna, and it just moves further into confusing territory when it comes to naming. The naming convention is especially difficult to decipher depending on what your native language is. Of course a latin language speaker might be able to easily determine oh yeah each one is slightly bigger than the other but I still think it borderlines too confus…
Based on the Intelligence vs. Cost graph, not clear to me why anyone would use Terra? Luna looks quite interesting though, happy to see OpenAI still serving the more budget-oriented side of the market (seems like Anthropic and Google have lost interest there). https://artificialanalysis.ai/articles/gpt-5-6-has-landed
- Better than Opus4.8 in the coding agent index (doubt)
- Just below sonnet 5, even with glm5.2, in the overall intelligence index
- Cheaper than haiku4.5, glm5.2 and kimi2.6 on cost per intelligence task index
We Openly hate OpenAI because they’re not very Open but we secretly hope they win against not-open-at-all Anthropic.