Live data from Hacker News

ARC-AGI-3

arcprize.org

131–140 of 394 posts

Re: ARC-AGI-3

#131
post #57

Some of these tasks are crazy. Even I can't beat them: https://arcprize.org/tasks/ar25

Just finished it, 8/8. I mostly approached it by winging it and shuffling things around that looked good and like it was approaching the goal, since there's plenty of time to finish. I still don't quite understand the exact mirroring rules at play.

I got stuck on 7/8 for a good while because I learned the rules wrong. I thought every bracket square needed to be lit.

Re: ARC-AGI-3

#132
post #89

Earlier quoted context omitted.

Well, yes, and would hand even more of an advantage to humans. My point is that designing a test around human advantages seems odd and orthogonal to measuring AGI.

The whole point of AGI is "general" intelligence, and for that intelligence to be broadly useful it needs to exist within the context of a human centric world

Then why deny it a harness it can also use in a human centric world?

Re: ARC-AGI-3

#133

Earlier quoted context omitted.

The main frontier models are all up on https://arcprize.org/tasks Barely any of them break 0% on any of the demo tasks, with Claude Opus 4.6 coming out on top with a few <3% scores, Gemini 3.1 Pro getting two nonzero scores, and the others (GPT-5.4 and Grok 4.20) getting all 0%

Curious, that doesn't match the graph up on the Leaderboard page? https://arcprize.org/leaderboard

The individual task scores are all on public tasks, they still held out a hundred or so private tasks that presumably GPT-5.4 did well on to get its leaderboard position.

Re: ARC-AGI-3

#134

Earlier quoted context omitted.

> Lastly, humans use way less energy to solve these in fewer steps, Not if you count all the energy that was necessary to feed, shelter and keep the the human at his preferred temperature so that he can sit in front of a computer and solve the problem.

ok, but thats the same for bulding a data center. Try again.

Yes, especially when considering a dataceter needed the energy of pretty many people to be built.

A single human is indeed more efficent, and way more flexible and actually just general intelligence.

Re: ARC-AGI-3

#135

https://x.com/scaling01 has called out a lot of issues with ARC-AGI-3, some of them (directly copied from tweets, with minimal editing): - Human baseline is "defined as the second-best first-run human by action count". Your "regular people" are people who signed up for puzzle solving and you don't compare the score against a human average but against the second best human solution - The scoring doesn't tell you how m…

"Very simplistic prompt" is the absolute and total core of this and the thing that ensures validity of the whole exercise.

If you are trying to measure GENERAL intelligence then it needs to be general.

Re: ARC-AGI-3

#136

https://x.com/scaling01 has called out a lot of issues with ARC-AGI-3, some of them (directly copied from tweets, with minimal editing): - Human baseline is "defined as the second-best first-run human by action count". Your "regular people" are people who signed up for puzzle solving and you don't compare the score against a human average but against the second best human solution - The scoring doesn't tell you how m…

If anything this makes the test much harder for the LLM to get high scores and that makes the scores they’re getting all that much more impressive.

Re: ARC-AGI-3

#137
post #11

My takeaway from playing a number of levels is that I am definitely not AGI

SGI - Sub General Intelligence or another more colloquial word commonly seen amongst users of wallstreetbets.

Re: ARC-AGI-3

#139
post #39

> As long as there is a gap between AI and human learning, we do not have AGI. Back in the 90's, Scientific American had an article on AI - I believe this was around the time Deep Blue beat Kasparov at chess. One AI researcher's quote stood out to me: "It's silly to say airplanes don't fly because they don't flap their wings the way birds do." He was saying this with regards to the Turing test, but I think the sentim…

> As long as there is a gap between AI and human learning, we do not have AGI.

Don't read the statement as a human dunk on LLMs, or even as philosophy.

The gap is important because of its special and devastating economic consequences. When the gap becomes truly zero, all human knowledge work is replaceable. From there, with robots, its a short step to all work is replaceable.

What's worse, the condition is sufficient but not even necessary. Just as planes can fly without flapping, the economy can be destroyed without full AGI.

Post reply on HN