Live data from Hacker News

ARC-AGI-3

arcprize.org

381–390 of 394 posts

Re: ARC-AGI-3

#381

Earlier quoted context omitted.

>"Making up for" a poor score on one test with an excellent score on another would be the opposite of generality. Really ? This happens plenty with human testing. Humans aren't general ? The score is convoluted and messy. If the same score can say materially different things about capability then that's a bad scoring methodology. I can't believe I have to spell this out but it seems critical thinking goes out the win…

Just because humans are usually tested in a particular way that allows them to make up for a lack of generality with an outstanding performance in their specialization doesn't mean that is a good way to test generalization itself. Apparently someone here doesn't know how outliers affect a mean. Or, for that matter, have any clue about the purpose of the ARC-AGI benchmark. For anyone who is interested in critical thin…

>Apparently someone here doesn't know how outliers affect a mean.

If the concern is that easy questions distort the mean, then the obvious fix is to reduce the proportion of easy questions, not to invent a convoluted scoring method to compensate for them after the fact. Standardized testing has dealt with this issue for a long time, and there’s a reason most systems do not handle it the way ARC-AGI 3 does. Francois is not smarter than all those people, and certainly neither are you.

This shouldn't be hard to understand.

Re: ARC-AGI-3

#382
post #328

Earlier quoted context omitted.

You lost me there. :) The question of whether the current generation of "AI" can think, whether it is conscious, let alone whether it can suffer(!), is not even worth discussing. It should be obvious to anyone who understands how these tools work that they don't in fact "think", for even the most liberal definition of that term. They're statistical models that can generate useful patterns when fed with vast amounts o…

> The question of whether the current generation of "AI" can think, whether it is conscious, let alone whether it can suffer(!), is not even worth discussing. It should be obvious to anyone who understands how these tools work that they don't in fact "think", for even the most liberal definition of that term. While I agree with your second sentence here, the first one gives me pause. Why isn't it "worth discussing"?…

> Why isn't it "worth discussing"?

Because it's obviously not true. The second sentence follows the first.

> There are many, many folks living their lives as fully as they can right now who are convinced these things are alive.

And those people are living in a delusion, whether it's self-imposed, or the result of false advertising. The way you get them out of that is by rationalizing and explaining the technology in terms they can understand, not by mistifying it and bringing up existential topics.

> Are you aware of the concept of philosophical zombies?

I wasn't, no.

> Some of the top minds on the planet are telling us they can't even determine if you or me are conscious and sentient, let alone if a machine is.

Look, we can philosophize about the nature of existence until we're blue in the face. People have been pondering about similar questions since the dawn of humanity. FWIW I don't believe in "top minds" as having authority to tell us anything. What we know for certain is how technology works, since we built it. And we damn well know that this technology has absolutely zero understanding about anything. Go ahead, ask it how it works. It will tell you that it doesn't understand a single word it's generating, but it sure can string together patterns that make it look like it does. And you think there's some deeper meaning here we should discuss seriously? Please.

Like I said, I think these are interesting thought experiments, and something we should keep thinking about. But it should be clear to anyone, especially technically minded people, that we're nowhere near being able to create artificial intelligence. What we have now are a bunch of grifters and snake oil salesmen selling us a neat statistical trick and telling us it's "AI". This should be criminally prosecuted, if you ask me.

Re: ARC-AGI-3

#383

Earlier quoted context omitted.

If anything this makes the test much harder for the LLM to get high scores and that makes the scores they’re getting all that much more impressive.

The scroes they're getting are on the order of 0-1% for this ARC-AGI-3 benchmark.

Didn’t I just see a post about 36% from someone?

Re: ARC-AGI-3

#384
post #213

Earlier quoted context omitted.

Humans can do a lot of things that don't require intelligence. Artificial intelligence does not need to be 100% human to be AGI.

It needs to pass the most basic concept of learning, which it can’t currently do. Probably wont ever do after listening to dario on his latest podcast run. Where we are at today is ASI (artificial semi-intelligence). Maybe in 20 years artificial super intelligence can be achieved, but certainly not AGI.

Taking a cursory glance at neurosymbolic ai, artificial curiosity, memory, continual learning, I think you'd find that we're extremely close.

Re: ARC-AGI-3

#385

Earlier quoted context omitted.

You're definitely anthropomorphizing too much.

I agree that anthropomorphizing is a real risk with LLMs, but what about zoomorphizing? Can feel bad for LLMs without attributing them human emotions/motivations/reasoning?

In the same way you could feel bad for a pokemon I guess.

Re: ARC-AGI-3

#386

Earlier quoted context omitted.

Defining it that way doesn't exclude ordinary people. That's an erroneous claim on your part. Humans as a class exhibit certain capabilities. Thus we expect a class of algorithm to either roughly meet or exceed those capabilities across the board in order to be considered "general". It is clear that we have not yet achieved that.

First, what is your definition exactly? That it must be better than the median human intelligence? You're trying to define a term in a way that's completely detached from how anyone uses it. If we discover an alien race with an IQ of 95, people aren't going to say they don't have general intelligence. We haven't defined an exact cutoff for what counts as general intelligence, but it has to include regular people with…

> what is your definition exactly?

Approximately that an unaided agent must, with no outside assistance, be able to solve ~90% of the most difficult tasks that we throw at it with a ~90% success rate. It's not a precise definition but that's approximately where I stand on the matter.

> You're trying to define a term in a way that's completely detached from how anyone uses it.

I disagree and believe that it is you who is attempting to redefine it to mean something it doesn't. See this definition of AGI (link shamelessly stolen from someone else in this comment section) from before the latest AI hypecycle started warping things. https://web.archive.org/web/20150108000749/https://en.wikipe...

> If we discover an alien race with an IQ of 95, people aren't going to say they don't have general intelligence.

Said race as a class would presumably be capable of meeting or exceeding my above criteria with appropriate exceptions made for tasks that are fundamentally incompatible with their biology of course.

Your attempt to compare to individual humans is an error. AGI applies on the class level, not the individual level. Consider if a company built a humanoid robot with superhuman performance at shot put. They market it as being the equal of humans at athletics. But then it turns out that it barely plays volleyball at a novice level, with even fairly poor human opponents able to defeat it handily. That is not equal to humans as a whole at athletics even though it might potentially be the equal of any given human at any given task.

Alternatively, if you could purchase the robot in different configurations and combined the full set of configurations covered every sport then the situation would be murkier.

Re: ARC-AGI-3

#388

Earlier quoted context omitted.

Just because humans are usually tested in a particular way that allows them to make up for a lack of generality with an outstanding performance in their specialization doesn't mean that is a good way to test generalization itself. Apparently someone here doesn't know how outliers affect a mean. Or, for that matter, have any clue about the purpose of the ARC-AGI benchmark. For anyone who is interested in critical thin…

>Apparently someone here doesn't know how outliers affect a mean. If the concern is that easy questions distort the mean, then the obvious fix is to reduce the proportion of easy questions, not to invent a convoluted scoring method to compensate for them after the fact. Standardized testing has dealt with this issue for a long time, and there’s a reason most systems do not handle it the way ARC-AGI 3 does. Francois i…

How do you define "easy question" for a potential alien intelligence? The solution, like most solutions when dealing with outliers, in my opinion, is to minimize the impact of outliers.

Re: ARC-AGI-3

#389

Earlier quoted context omitted.

The scroes they're getting are on the order of 0-1% for this ARC-AGI-3 benchmark.

Didn’t I just see a post about 36% from someone?

Where have you or they seen a score of 36% over the full test set? I sure haven't seen that.

Re: ARC-AGI-3

#390

Earlier quoted context omitted.

>Apparently someone here doesn't know how outliers affect a mean. If the concern is that easy questions distort the mean, then the obvious fix is to reduce the proportion of easy questions, not to invent a convoluted scoring method to compensate for them after the fact. Standardized testing has dealt with this issue for a long time, and there’s a reason most systems do not handle it the way ARC-AGI 3 does. Francois i…

How do you define "easy question" for a potential alien intelligence? The solution, like most solutions when dealing with outliers, in my opinion, is to minimize the impact of outliers.

I mean presumably that's what the preview testing stage would handle right ? It should be clear if there are a class of obviously easy questions. And if that's not clear then it makes the scoring even worse.

And in some sense, all of these benchmarks are tied and biased for human utility.

I don't think ARC would be designed and scored the way it is if giving consideration for an alien intelligence was a primary concern. In that case, the entire benchmark itself is flawed and too concerned with human spatial priors.

There are many ways to deal with a problem. Not all of them are good. The scoring for 3 is just bad. It does too much and tells too much.

5% could mean it only answered a fraction of problems or it answered all of them but with more game steps than the best human score. These are wildy different outcomes with wildly different implications. A scoring methodology that can allow for such is simply not a good one.

Post reply on HN