GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game https://arcprize.org/results/openai-gpt-5-6
This is the first I have herd of this benchmark. Can someone explain how it in any way indicates how close we are to "AGI"? Replay of Sol attempting the game: https://arcprize.org/replay/83543d22-8e1e-439a-8809-129ff1d9... It seems a weird and arbitrary challenge for a language model to be expected to perform. It also seems like there are some harness/visual issues even in the first few steps, where it states that it…
GPT-5.6
761–770 of 1001 posts
Re: GPT-5.6
#762I really wish there was just an easy guide on when to use Sol vs Terra vs Luna, and it just moves further into confusing territory when it comes to naming. The naming convention is especially difficult to decipher depending on what your native language is. Of course a latin language speaker might be able to easily determine oh yeah each one is slightly bigger than the other but I still think it borderlines too confus…
I use the strongest model (5.5 now 5.6 sol) on the highest reasoning effort with /fast for everything. With a $200 pro sub I can't even use my weekly limit. And it's faster than using a weaker model that makes more mistakes which I have to waste time fixing.
Re: GPT-5.6
#763We Openly hate OpenAI because they’re not very Open but we secretly hope they win against not-open-at-all Anthropic.
Re: GPT-5.6
#764Earlier quoted context omitted.
Very interesting. My prediction is that Mythos would outperform Sol. Also what does this tell about Yann LeCuns whole world model theory? Bro has been going on and on about it. He has made multiple wrong predictions on the trajectory of LLMs. At some point his claim should be fully falsified no?
Mythos doesn't appear to be on the verified leaderboard for ARC-AGI 3
Re: GPT-5.6
#765GPT-5.6 is a really good model, and quite cheap. I can finally replace GPT-5.3-Codex for my Tool Calling in n8n. Here's my benchmark results for GPT-5.6: https://aibenchy.com/?q=gpt-5.6 (the high reasoning variants are still running, uploading them soon too) EDIT: The high variants are there too, enjoy the hamsters[0]. [0]: https://aibenchy.com/showcase/?q=gpt-5.6
[flagged]
Re: GPT-5.6
#766Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?
Not having to deal with Anthropics constantly changing policies, token-gating, and carrot-and-stick marketing helps me to focus on work, rather than dealing with their company problems.
Re: GPT-5.6
#767I love testing the new models by asking them to code a toy RTS game. Here's what Terra did: https://senko.net/vibecode-bench/2026/rts-gpt-5.6-terra.html (one try, in codex app, xhigh effort) Comparing this to other models, I find it similar to GPT-5.5 and a bit behind Sonnet 5. You can see how other models fared here: https://senko.net/vibecode-bench/ (you can also fetch the prompt and the the 5.6 Terra resulting cod…
Re: GPT-5.6
#768Re: GPT-5.6
#769Earlier quoted context omitted.
You can't post like this to HN, regardless of how you feel about someone else's posts. Doing it repeatedly crosses into harassment, and you've done it more than 3 times now - e.g. https://news.ycombinator.com/item?id=48504823 https://news.ycombinator.com/item?id=48504654 So please especially stop doing that. If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit…
Those 2 comments are from a month ago, on the same day , and my point was he was spamming those threads with basically the same comment, and he links often to his site and his posts are all about his software. Both of my replies you link have positive counts (with sure had plenty of downvotes as he had positive upvotes), and even now after your link probably added more downvotes to mine. And other commenters in those…
Re: GPT-5.6
#770Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?
I sub both codex and claude at 20x. I like opus+fable more than gpt5.5 because it seems gpt tries to finish tasks by leaving any ambiguity unresolved. claude seems better at surfacing open questions. This is using the same AGENTS.md prompts, which were designed firstly for Claude use, so maybe it's something that could be optimized better if I understood gpt as well?