I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any…
Does it experience?
Can it connect with other agents, understand them, come to empathise with them and find a way to work with them better?
The answer is no to all of these, and there are other problems as well. Yes, this model is trained to use a domain specific language to reason and plan over puzzle problems, and so it's programmers have cracked arc-agi-3 and that's a great achievement, but there is an asymmetry here. The arc team are well funded but are charged with providing a target for the vast ocean of funding, compute and talent everywhere else.
Most importantly, arc-agi-3 and the other benchmarks are all verifiable. The model can check if it's succeeded or not. They are not A* of course, but long horizon problems where you have to overcome minima to get the solution are not alien to AI either.