Live data from Hacker News

GPT-5.6

openai.com

761–770 of 1001 posts

Re: GPT-5.6

#761

GPT-5.6 Sol sets a new SOTA on ARC-AGI-3: 7.8% Sol is the first verified frontier model to ever beat an ARC-AGI-3 game https://arcprize.org/results/openai-gpt-5-6

This is the first I have herd of this benchmark. Can someone explain how it in any way indicates how close we are to "AGI"? Replay of Sol attempting the game: https://arcprize.org/replay/83543d22-8e1e-439a-8809-129ff1d9... It seems a weird and arbitrary challenge for a language model to be expected to perform. It also seems like there are some harness/visual issues even in the first few steps, where it states that it…

the problems are general and abstract, domain specific knowledge and memorization don't help. Figuring out the rules, the goal, the controls, and how to solve in a reasonable budget all indicate some level of general ability.

Re: GPT-5.6

#762

I really wish there was just an easy guide on when to use Sol vs Terra vs Luna, and it just moves further into confusing territory when it comes to naming. The naming convention is especially difficult to decipher depending on what your native language is. Of course a latin language speaker might be able to easily determine oh yeah each one is slightly bigger than the other but I still think it borderlines too confus…

I use the strongest model (5.5 now 5.6 sol) on the highest reasoning effort with /fast for everything. With a $200 pro sub I can't even use my weekly limit. And it's faster than using a weaker model that makes more mistakes which I have to waste time fixing.

I have found that higher reasoning doesn't always produce better results. Models often tend to overthink and get themselves in a bind. And tasks just take longer for no reason.

Re: GPT-5.6

#763

We Openly hate OpenAI because they’re not very Open but we secretly hope they win against not-open-at-all Anthropic.

I'm just happy there's competition.

Re: GPT-5.6

#764

Earlier quoted context omitted.

Very interesting. My prediction is that Mythos would outperform Sol. Also what does this tell about Yann LeCuns whole world model theory? Bro has been going on and on about it. He has made multiple wrong predictions on the trajectory of LLMs. At some point his claim should be fully falsified no?

Mythos doesn't appear to be on the verified leaderboard for ARC-AGI 3

Officially, Fable/Mythos testing was delayed because of Anthropic's data retention policy. Don't know if there's word on them working that out yet.

https://x.com/arcprize/status/2064399134099153344

Re: GPT-5.6

#765
post #597

GPT-5.6 is a really good model, and quite cheap. I can finally replace GPT-5.3-Codex for my Tool Calling in n8n. Here's my benchmark results for GPT-5.6: https://aibenchy.com/?q=gpt-5.6 (the high reasoning variants are still running, uploading them soon too) EDIT: The high variants are there too, enjoy the hamsters[0]. [0]: https://aibenchy.com/showcase/?q=gpt-5.6

[flagged]

Where's your website?

Re: GPT-5.6

#766
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

I used to have the CC $200 plan, and moved to Codex 6 months ago. I have the anthropic $20 plan + API billing for rare use. Use Codex daily.

Not having to deal with Anthropics constantly changing policies, token-gating, and carrot-and-stick marketing helps me to focus on work, rather than dealing with their company problems.

Re: GPT-5.6

#767
post #439

I love testing the new models by asking them to code a toy RTS game. Here's what Terra did: https://senko.net/vibecode-bench/2026/rts-gpt-5.6-terra.html (one try, in codex app, xhigh effort) Comparing this to other models, I find it similar to GPT-5.5 and a bit behind Sonnet 5. You can see how other models fared here: https://senko.net/vibecode-bench/ (you can also fetch the prompt and the the 5.6 Terra resulting cod…

I've been wanting to make an RTS but found it a bit daunting. I thought this might be a bit challenging for LLMs too but I guess not! This is nice to see.

Re: GPT-5.6

#768
It scores high on BenchCAD, that's interesting to see, I was wondering about how each model could handle this. Seems like they trained it on programmatic CAD specifically.

Re: GPT-5.6

#769
post #662
post #541

Earlier quoted context omitted.

You can't post like this to HN, regardless of how you feel about someone else's posts. Doing it repeatedly crosses into harassment, and you've done it more than 3 times now - e.g. https://news.ycombinator.com/item?id=48504823 https://news.ycombinator.com/item?id=48504654 So please especially stop doing that. If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit…

Those 2 comments are from a month ago, on the same day , and my point was he was spamming those threads with basically the same comment, and he links often to his site and his posts are all about his software. Both of my replies you link have positive counts (with sure had plenty of downvotes as he had positive upvotes), and even now after your link probably added more downvotes to mine. And other commenters in those…

Why do you care so much about his posts? It's not a bad benchmark, it clearly isn't saturated, and allows models to be vibe-compared over time.

Re: GPT-5.6

#770
post #6

Ok long time Claude Code user here; lately I've started to realize there's other great models out there I should be trying, but I'm hesitant to leave Claude Code behind for something new. What's the consensus today on codex vs claude code, does it really matter anymore?

I sub both codex and claude at 20x. I like opus+fable more than gpt5.5 because it seems gpt tries to finish tasks by leaving any ambiguity unresolved. claude seems better at surfacing open questions. This is using the same AGENTS.md prompts, which were designed firstly for Claude use, so maybe it's something that could be optimized better if I understood gpt as well?

Is it you have Fable delegate work to Sol? How do you do that? Do you run it in Codex/Desktop app?
Post reply on HN