Live data from Hacker News

ARC-AGI-3

arcprize.org

21–30 of 394 posts

Re: ARC-AGI-3

#21
post #12

Any benchmarks?

The main frontier models are all up on https://arcprize.org/tasks

Barely any of them break 0% on any of the demo tasks, with Claude Opus 4.6 coming out on top with a few <3% scores, Gemini 3.1 Pro getting two nonzero scores, and the others (GPT-5.4 and Grok 4.20) getting all 0%

Re: ARC-AGI-3

#23

what is the evidence that being able to play games equates to AGI?

The evidence is that humans are able to win these games. AGI is usually defined as the ability to do any intellectual task about as well as a highly competent human could. The point of these ARC benchmarks is to find tasks that humans can do easily and AI cannot, thus driving a new reasoning competency as companies race each other to beat human performance on the benchmark.

Re: ARC-AGI-3

#24

Earlier quoted context omitted.

This is clear AGI progress. It should show you, that AI is not sleeping, it gets better and you should use this as a signal that you should take this topic serious.

Labelling a test "AGI" does not show AGI progress any more than labelling a cpu "AGI" makes it so. It might show that AI tools are improving but it does not necessarily follow that tools improving = AGI progress if you're on the completely wrong trail.

The transfer of knowledge required here is that a ARC-AGI-3 is now necessary and adds another dimension of capability.

These 'tests' are not labeled AGI by magic but because they are designed specificly for testing certain things a question answer test cant solve.

Gemini and OpenAI are at 80-90% at ARC-AGI-2 and its quite interesting to see the difference of challange between 2 and 3.

AGI progress means btw. general. So every additional dimension an agent can solve pushes that agent to be more general.

Re: ARC-AGI-3

#25

what is the evidence that being able to play games equates to AGI?

I think the idea is that if they cannot perform any cognitive task that is trivial for humans then we can state they haven’t reached ‘AGI’.

It used to be easy to build these tests. I suspect it’s getting harder and harder.

But if we run out of ideas for tests that are easy for humans but impossible for models, it doesn’t mean none exist. Perhaps that’s when we turn to models to design candidate tests, and have humans be the subjects to try them out ad nauseam until no more are ever uncovered? That sounds like a lovely future…

Re: ARC-AGI-3

#26

Earlier quoted context omitted.

This is clear AGI progress. It should show you, that AI is not sleeping, it gets better and you should use this as a signal that you should take this topic serious.

Labelling a test "AGI" does not show AGI progress any more than labelling a cpu "AGI" makes it so. It might show that AI tools are improving but it does not necessarily follow that tools improving = AGI progress if you're on the completely wrong trail.

[deleted]

Re: ARC-AGI-3

#28

what is the evidence that being able to play games equates to AGI?

There isn't a strict definition of AGI, there's no way to find evidence for what equates to it, and besides, things like this are meant only as likely necessary conditions. Anyway, from the article: > As long as there is a gap between AI and human learning, we do not have AGI. This seems like a reasonable requirement. Something I think about a lot with vibe coding is that unlike humans, individual models do not get b…

Is that within a codebase off relatively fixed size that things get worse as time goes on, or are you saying as the codebase grows that the limits of a model's context means that because the model is no longer able to hold the entire codebase within its context that it performs worse than when the codebase was smaller?
Post reply on HN