Live data from Hacker News

GPT-5.6

openai.com

671–680 of 1001 posts

Re: GPT-5.6

#673

I use both Claude and Codex, but mostly Claude for planning and coding, and Codex to review Claude’s work. I follow a sort of waterfall workflow which is verbose but fully transparent. Anthropic’s $100 subscription works fine for me, but whatever subscription my company has with OpenAI reaches the 5hr limit ridiculously quickly.

When I was going through this it was because OpenAI had defaulted to /fast mode with 2x token usage

Re: GPT-5.6

#675

GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game https://arcprize.org/results/openai-gpt-5-6

This is the first I have herd of this benchmark. Can someone explain how it in any way indicates how close we are to "AGI"?

Replay of Sol attempting the game: https://arcprize.org/replay/83543d22-8e1e-439a-8809-129ff1d9...

It seems a weird and arbitrary challenge for a language model to be expected to perform. It also seems like there are some harness/visual issues even in the first few steps, where it states that it hasn't moved when it clearly has.

Re: GPT-5.6

#676
post #379

Here are 18 pelicans - six each for Luna, Terra and Sol at the six different reasoning effort levels (plus the price to generate each one): https://static.simonwillison.net/static/2026/gpt-5.6-pelican... Or if you want to see some in 3D, OpenAI featured a pelican riding a tricycle, bicycle, pony and another pelican in their livestream this morning: https://www.youtube.com/live/Wq45rvPGNHs?t=1070s

Time to dump this test. Probably not a coincidence every version has the same rolling green hills, gradient blue sky, sun in the corner, etc.

I partially agree, but in this case it kinda illustrates that it may not be worth using Terra on any reasoning level below high; those are some awful penguins on bikes.

Re: GPT-5.6

#677

I really wish there was just an easy guide on when to use Sol vs Terra vs Luna, and it just moves further into confusing territory when it comes to naming. The naming convention is especially difficult to decipher depending on what your native language is. Of course a latin language speaker might be able to easily determine oh yeah each one is slightly bigger than the other but I still think it borderlines too confus…

I use the strongest model (5.5 now 5.6 sol) on the highest reasoning effort with /fast for everything. With a $200 pro sub I can't even use my weekly limit. And it's faster than using a weaker model that makes more mistakes which I have to waste time fixing.

Re: GPT-5.6

#678
post #503

Based on the Intelligence vs. Cost graph, not clear to me why anyone would use Terra? Luna looks quite interesting though, happy to see OpenAI still serving the more budget-oriented side of the market (seems like Anthropic and Google have lost interest there). https://artificialanalysis.ai/articles/gpt-5-6-has-landed

Luna@max is in a VERY interesting spot if their rankings are at all to be believed:

- Better than Opus4.8 in the coding agent index (doubt)

- Just below sonnet 5, even with glm5.2, in the overall intelligence index

- Cheaper than haiku4.5, glm5.2 and kimi2.6 on cost per intelligence task index

Re: GPT-5.6

#680
5.6 SOL is basically useless, even on fast mode. It takes so long to do anything that it would be faster to do yourself. And it burns usage so quickly it's genuinely not worth it.
Post reply on HN