Some of these tasks are crazy. Even I can't beat them: https://arcprize.org/tasks/ar25
Just finished it, 8/8. I mostly approached it by winging it and shuffling things around that looked good and like it was approaching the goal, since there's plenty of time to finish. I still don't quite understand the exact mirroring rules at play.
ARC-AGI-3
131–140 of 394 posts
Re: ARC-AGI-3
#132Earlier quoted context omitted.
Well, yes, and would hand even more of an advantage to humans. My point is that designing a test around human advantages seems odd and orthogonal to measuring AGI.
The whole point of AGI is "general" intelligence, and for that intelligence to be broadly useful it needs to exist within the context of a human centric world
Re: ARC-AGI-3
#133Earlier quoted context omitted.
The main frontier models are all up on https://arcprize.org/tasks Barely any of them break 0% on any of the demo tasks, with Claude Opus 4.6 coming out on top with a few <3% scores, Gemini 3.1 Pro getting two nonzero scores, and the others (GPT-5.4 and Grok 4.20) getting all 0%
Curious, that doesn't match the graph up on the Leaderboard page? https://arcprize.org/leaderboard
Re: ARC-AGI-3
#134Earlier quoted context omitted.
> Lastly, humans use way less energy to solve these in fewer steps, Not if you count all the energy that was necessary to feed, shelter and keep the the human at his preferred temperature so that he can sit in front of a computer and solve the problem.
ok, but thats the same for bulding a data center. Try again.
A single human is indeed more efficent, and way more flexible and actually just general intelligence.
Re: ARC-AGI-3
#135https://x.com/scaling01 has called out a lot of issues with ARC-AGI-3, some of them (directly copied from tweets, with minimal editing): - Human baseline is "defined as the second-best first-run human by action count". Your "regular people" are people who signed up for puzzle solving and you don't compare the score against a human average but against the second best human solution - The scoring doesn't tell you how m…
If you are trying to measure GENERAL intelligence then it needs to be general.
Re: ARC-AGI-3
#136https://x.com/scaling01 has called out a lot of issues with ARC-AGI-3, some of them (directly copied from tweets, with minimal editing): - Human baseline is "defined as the second-best first-run human by action count". Your "regular people" are people who signed up for puzzle solving and you don't compare the score against a human average but against the second best human solution - The scoring doesn't tell you how m…
Re: ARC-AGI-3
#137My takeaway from playing a number of levels is that I am definitely not AGI
Re: ARC-AGI-3
#138At this point, I'm pretty sure we'll just know when it happens.
Re: ARC-AGI-3
#139> As long as there is a gap between AI and human learning, we do not have AGI. Back in the 90's, Scientific American had an article on AI - I believe this was around the time Deep Blue beat Kasparov at chess. One AI researcher's quote stood out to me: "It's silly to say airplanes don't fly because they don't flap their wings the way birds do." He was saying this with regards to the Turing test, but I think the sentim…
Don't read the statement as a human dunk on LLMs, or even as philosophy.
The gap is important because of its special and devastating economic consequences. When the gap becomes truly zero, all human knowledge work is replaceable. From there, with robots, its a short step to all work is replaceable.
What's worse, the condition is sufficient but not even necessary. Just as planes can fly without flapping, the economy can be destroyed without full AGI.