Live data from Hacker News

Arc-AGI-2 and ARC Prize 2025

arcprize.org

21–30 of 103 posts

Re: Arc-AGI-2 and ARC Prize 2025

#21
Have you had any neurologists utilize your dataset? My own reaction after solving a few of the puzzles was "Why is this so intuitive for me, but not for an LLM?".

Our human-ability to abstract things is underrated.

Re: Arc-AGI-2 and ARC Prize 2025

#22
> and was the only benchmark to pinpoint the exact moment in late 2024 when AI moved beyond pure memorization

This is self-referential, the benchmark pinpointed the time when AI went from memorization to problem solving, because the benchmark requires problem solving to complete. How do we know it requires problem solving skills? Because memorization-only LLMs can't do it but humans can.

I think ARC are producing some great benchmarks, and I think they probably are pushing forward the state of the art, however I don't think they identified anything particular with o3, at least they don't seem to have proven a step change.

Re: Arc-AGI-2 and ARC Prize 2025

#23
post #6

Earlier quoted context omitted.

ARC 3 is still spatially 2D, but it adds a time dimension, and it's interactive.

I think a lot of people got discouraged, seeing how openai solved arc agi 1 by what seems like brute forcing and throwing money at it. Do you believe arc was solved in the "spirit" of the challenge? Also all the open sourced solutions seem super specific to solving arc. Is this really leading us to human level AI at open ended tasks?

It's useful to know what current AI systems can achieve with unlimited test-time compute resources. Ultimately though, the "spirit of the challenge" is efficiency, which is why we're specifically looking for solutions that are at least within 1-2 order of magnitude of cost from being competitive with humans. The Kaggle leaderboard is very resource-constrained, and on the public leaderboard you need to use less than $10,000 in compute to solve 120 tasks.

Re: Arc-AGI-2 and ARC Prize 2025

#24
post #21

Have you had any neurologists utilize your dataset? My own reaction after solving a few of the puzzles was "Why is this so intuitive for me, but not for an LLM?". Our human-ability to abstract things is underrated.

There have been some human studies on ARC 1 previously, I expect there will be more in the future. See this paper from 2021, which was one of the earliest works in this direction: https://arxiv.org/abs/2103.05823

Re: Arc-AGI-2 and ARC Prize 2025

#25
I'd very much like to see VLAs get in the game with ARC. When I solve these puzzles I'm imagining myself move blocks around. Much of the time I'm treating these as physics simulations with custom physics per puzzle. VLAs are particularly well suited to the kind of training and planning which might unlock solutions here.

Re: Arc-AGI-2 and ARC Prize 2025

#26
post #6

Oh boy! Some of these tasks are not hard, but require full attention and a lot of counting just to get things right! ARC3 will go 3D perhaps? JK Congrats on launch, lets see how long it'll take to get saturated

ARC 3 is still spatially 2D, but it adds a time dimension, and it's interactive.

If you aren't joking, that will filter most humans.

Re: Arc-AGI-2 and ARC Prize 2025

#27

> and was the only benchmark to pinpoint the exact moment in late 2024 when AI moved beyond pure memorization This is self-referential, the benchmark pinpointed the time when AI went from memorization to problem solving, because the benchmark requires problem solving to complete. How do we know it requires problem solving skills? Because memorization-only LLMs can't do it but humans can. I think ARC are producing som…

The reason these tasks require fluid intelligence is because they were designed this way -- with task uniqueness/novelty as the primary goal.

ARC 1 was released long before in-context learning was identified in LLMs (and designed before Transformer-based LLMs existed), so the fact that LLMs can't do ARC was never a design consideration. It just turned out this way, which confirmed our initial assumption.

Re: Arc-AGI-2 and ARC Prize 2025

#28

Oh boy! Some of these tasks are not hard, but require full attention and a lot of counting just to get things right! ARC3 will go 3D perhaps? JK Congrats on launch, lets see how long it'll take to get saturated

The "select" tool gives some help with tasks that require counting or copying. You can select areas of the input, which will show their dimensions, and copy-paste them into the output (ctrl+c/ctrl+v).

Re: Arc-AGI-2 and ARC Prize 2025

#29
post #6

Oh boy! Some of these tasks are not hard, but require full attention and a lot of counting just to get things right! ARC3 will go 3D perhaps? JK Congrats on launch, lets see how long it'll take to get saturated

ARC 3 is still spatially 2D, but it adds a time dimension, and it's interactive.

Are you in the process of creating tasks that behave as an acid test for AGI? If not, do you think such a task is feasible? I read somewhere in the ARC blog that they define AGI as when creating tasks that is hard for AI but easy for humans becomes virtually impossible.

Re: Arc-AGI-2 and ARC Prize 2025

#30
Maybe this is a really stupid question but I've been curious... are LLMs based on... "Neuronormativity"? Like, what neurology is an LLM based on? Would we get any benefit from looking at neurodiverse processing styles?
Post reply on HN