Live data from Hacker News

ARC-AGI-3

arcprize.org

391–394 of 394 posts

Re: ARC-AGI-3

#391

Earlier quoted context omitted.

First, what is your definition exactly? That it must be better than the median human intelligence? You're trying to define a term in a way that's completely detached from how anyone uses it. If we discover an alien race with an IQ of 95, people aren't going to say they don't have general intelligence. We haven't defined an exact cutoff for what counts as general intelligence, but it has to include regular people with…

> what is your definition exactly? Approximately that an unaided agent must, with no outside assistance, be able to solve ~90% of the most difficult tasks that we throw at it with a ~90% success rate. It's not a precise definition but that's approximately where I stand on the matter. > You're trying to define a term in a way that's completely detached from how anyone uses it. I disagree and believe that it is you who…

> See this definition of AGI (link shamelessly stolen from someone else in this comment section) from before the latest AI hypecycle started warping things.

Every definition on that page, both theoretical and operational, match my definition and not yours. Notice that none of them would exclude an AGI with an IQ around 90, provided it's intelligence is general.

> Approximately that an unaided agent must, with no outside assistance, be able to solve ~90% of the most difficult tasks that we throw at it with a ~90% success rate. It's not a precise definition but that's approximately where I stand on the matter.

This isn't your definition. How hard are these "most difficult tasks"? Can 50% of humans solve them? 10%? If it were literally the most difficult problems, they would be the ones 0% of humans have solved.

> Said race as a class would presumably be capable of meeting or exceeding my above criteria

Some might, but this hypothetical alien race does not. Do you still consider them to have general intelligence if they can merely do everything a 95 IQ person could?

> Your attempt to compare to individual humans is an error. AGI applies on the class level

According only to you. LLMs are benchmarked individually. No one runs a benchmark where Claude gets half the questions right, GPT gets the other half right, and it's reported as a combined perfect score as a class. Instead the each score 50%. (Not that I think current AIs can solve the harder benchmark problems. The point is they are measured individually.)

No one else ascribes general intelligence only to a class. You can talk to one average person (or alien), give them some tests, and determine they have general intelligence. This is how everyone else uses the term.

Re: ARC-AGI-3

#393
post #310

Earlier quoted context omitted.

>Talking to the ARC folks tonight, it sounds like there will be an ARC-4,5,6,etc. I mean of course there will be. Quintessential goal post moving...

If you read the charter of the eval (or any eval, really), this statement is pretty silly. The whole point of each eval version is to identify a chunk of challenges that humans do well that AI can't. When AI gets to ~80, you move to the next chunk. When you run out of challenges, you have AGI.

Except you will never run out of challenges and my sense from Chollet has been that every challenge was hinted at being the final one where once beaten AGI would have been created and of course at the end of each one he comes out saying akshuallyyyy this isn't AGI and it wont be AGI until ARC Challenge+1 is beaten!

Re: ARC-AGI-3

#394

It's getting pretty old now when Francois Chollet puts out a new ARC challenge, claims definitively that no system is going to crack it without being full blown AGI, the benchmark gets saturated in a few months, he claims the systems definitely aren't AGI then puts out a new challenge that no non AGI system can clear and a few months later.... etc. etc.

Chollet literally never says that. Quite the opposite. He says that AIs are currently abysmally bad at the skills this benchmark tests. An AGI should be able to do this, but doing this doesn't mean it's AGI . He has been very clear about that. I suggest you go back and (re)read the intro ARC-AGI paper. No system can crack these out of the box (like humans can) because we don't have AGI.

Yeah I mean ChatGPT 5.4 Pro can't even pick my nose for me so it's obviously not AGI /s
Post reply on HN