Live data from Hacker News

ARC-AGI-3

arcprize.org

181–190 of 394 posts

Re: ARC-AGI-3

#181

Earlier quoted context omitted.

AGI’s 'general' is the wrong word, I thinkg. Humans aren’t general, we’re jagged. Strong in some areas, weak in others, and already surpassed in many domains. LLM are way past us at languages for instance. Calculators passed us at calculating, etc.

We are jagged, but we can smooth that jaggedness if we choose to do so. LLMs stay jagged.

There's no objective measure of intelligence comparisons, we only say llm is jagged compared to humans.

Re: ARC-AGI-3

#182

Earlier quoted context omitted.

If I had a puzzle I really needed solved, then I would not ask a rando on the street, I would ask someone I know is really good at puzzles. My point is: For AGI to be useful, it really should be able to perform at the top 10% or better level for as many professions as possible (ideally all of them). An AI that can only perform at the average human level is useless unless it can be trained for the job like humans can.

> An AI that can only perform at the average human level is useless unless it can be trained for the job like humans can. Yes, if you want skilled labour. But that's not at all what ARC-AGI attempts to test for: it's testing for general intelligence as possessed by anyone without a mental incapacity.

It seems they don't test for that, since they use the second-best human solution as a baseline.

And that's the right way to go. When computers were about to become superhuman at chess, few people cared that it could beat random people for many years prior to that. They cared when Kasparov was dethroned.

Remember, the point here is marketing as well as science. And the results speak for themselves. After all, you remember Deep Blue, and not the many runners-up that tried. The only reason you remember is because it beat Kasparov.

Re: ARC-AGI-3

#183
I hope at least some of these are direct Chip's Challenge ports. Waiting for some old muscle memory to kick in here.

Re: ARC-AGI-3

#184
post #132

Earlier quoted context omitted.

The whole point of AGI is "general" intelligence, and for that intelligence to be broadly useful it needs to exist within the context of a human centric world

Then why deny it a harness it can also use in a human centric world?

There is no general purpose harness.

Re: ARC-AGI-3

#185

Earlier quoted context omitted.

If I had a puzzle I really needed solved, then I would not ask a rando on the street, I would ask someone I know is really good at puzzles. My point is: For AGI to be useful, it really should be able to perform at the top 10% or better level for as many professions as possible (ideally all of them). An AI that can only perform at the average human level is useless unless it can be trained for the job like humans can.

> An AI that can only perform at the average human level is useless unless it can be trained for the job like humans can. Yes, if you want skilled labour. But that's not at all what ARC-AGI attempts to test for: it's testing for general intelligence as possessed by anyone without a mental incapacity.

[deleted]

Re: ARC-AGI-3

#186
I feel like we've got tunnel vision. Things you can do on a computer are a tiny subset of what a human can do.

If the AI has to control a body to sit on a couch and play this game on a laptop that would be a step in the right direction.

Re: ARC-AGI-3

#187
post #145

This is a very good estimation of AGI. We give humans and AI the same input and measure the results. Kudos to ARC for creating these games. I really wonder why so many people fight against this. We know that AI is useful, we know that AI is researchful, but we want to know if they are what we vaguely define as intelligence. I’ve read the airplanes don’t use wings, or submarines don’t swim. Yes, but this is is not the…

AGI’s 'general' is the wrong word, I thinkg. Humans aren’t general, we’re jagged. Strong in some areas, weak in others, and already surpassed in many domains. LLM are way past us at languages for instance. Calculators passed us at calculating, etc.

We don't call a calculator intelligent.

A calculator is extremely useful, but it is not intelligent.

A computer is extremely useful, but it is not intelligent.

Airplanes don't have wings, but they're damn sure useful, and also not intelligent.

If LLMs cannot learn to beat not-that-difficult of games better than young teens, they are not intelligent.

They are extremely useful. But they are not AGI.

Words matter.

Re: ARC-AGI-3

#188

Earlier quoted context omitted.

AGI’s 'general' is the wrong word, I thinkg. Humans aren’t general, we’re jagged. Strong in some areas, weak in others, and already surpassed in many domains. LLM are way past us at languages for instance. Calculators passed us at calculating, etc.

We don't call a calculator intelligent. A calculator is extremely useful, but it is not intelligent. A computer is extremely useful, but it is not intelligent. Airplanes don't have wings, but they're damn sure useful, and also not intelligent. If LLMs cannot learn to beat not-that-difficult of games better than young teens, they are not intelligent. They are extremely useful. But they are not AGI. Words matter.

> If LLMs cannot learn to beat not-that-difficult of games better than young teens, they are not intelligent.

I agree, with unresolved questions. Does it count if the LLM writes code which trains a neural network to play the game, and that neural network plays the game better than people do? Does that only count if the LLM tries that solution without a human prompting it to do so?

Re: ARC-AGI-3

#189

https://x.com/scaling01 has called out a lot of issues with ARC-AGI-3, some of them (directly copied from tweets, with minimal editing): - Human baseline is "defined as the second-best first-run human by action count". Your "regular people" are people who signed up for puzzle solving and you don't compare the score against a human average but against the second best human solution - The scoring doesn't tell you how m…

Those are supposed to be issues? After reading your list my impression of ARC-AGI has gone up rather than down. All of those things seem like the right way to go about this.

They are severe problems if your income is tied to LLM hype generation.

Re: ARC-AGI-3

#190

Earlier quoted context omitted.

>We also observed a case where a user created a loop that repeatedly called a model and asked for the time. Given the user role’s odd and repetitive behavior, the model could easily tell it was also controlled by an automated system of some kind. Over many iterations, the model began to exhibit “fed up” behavior and attempted to prompt-inject the system controlling the user role. The injection attempted to override p…

If this is a serious risk we should pull the plug now while we can still reach it. If we have to rely on the mood and temperament of LLMs for security, we're already lost.

Welcome to the ride, people have been talking about this for at least 15 years now.

I mean, the original plan that pretty much every one agreed on was to absolutely not give it access to the internet. Which already went out the window on day one.

Post reply on HN