Live data from Hacker News

François Chollet: The Arc Prize and How We Get to AGI [video]

youtube.com

141–150 of 230 posts

Re: François Chollet: The Arc Prize and How We Get to AGI [video]

#141
post #98

Earlier quoted context omitted.

Getting a high score on ARC doesn't mean we have AGI and Chollet has always said as much AFAIK, it's meant to push the AI research space in a positive direction. Being able to solve ARC problems is probably a pre-requisite to AGI. It's a directional push into the fog of war, with the claim being that we should explore that area because we expect it's relevant to building AGI.

"Being able to solve ARC problems is probably a pre-requisite to AGI." - is it? Humans have general intelligence and most can't solve the harder ARC problems.

https://arcprize.org/leaderboard

"Avg. Mturker" has 77% on ARC1 and costs $3/task. "Stem Grad" has 98% on ARC1 and costs $10/task. I would love a segment like "typical US office employee" or something else in between since I don't think you need a stem degree to do better than 77%.

It's also worth noting the "Human Panel" gets 100% on ARC2 at $17/task. All the "Human" models are on the score/cost frontier and exceptional in their score range although too expensive to win the prize obviously.

I think the real argument is that the ARC problems are too abstract and obscure to be relevant to useful AGI, but I think we need a little flexibility in that area so we can have tests that can be objectively and mechanically graded. E.g. "write a NYT bestseller" is an impractical test in many ways even if it's closer to what AGI should be.

Re: François Chollet: The Arc Prize and How We Get to AGI [video]

#142
post #4

I feel like I'm the only one who isn't convinced getting a high score on the ARC eval test means we have AGI. It's mostly about pattern matching (and some of it ambiguous even for humans what the actual true response aught to be). It's like how in humans there's lots of different 'types' of intelligence, and just overfitting on IQ tests doesn't in my mind convince me a person is actually that smart.

Getting a high score on ARC doesn't mean we have AGI and Chollet has always said as much AFAIK, it's meant to push the AI research space in a positive direction. Being able to solve ARC problems is probably a pre-requisite to AGI. It's a directional push into the fog of war, with the claim being that we should explore that area because we expect it's relevant to building AGI.

> Getting a high score on ARC doesn't mean we have AGI and Chollet has always said as much AFAIK

He only seems to say this recently, since OpenAI cracked the ARC-AGI benchmark. But in the original 2019 abstract he said this:

> We argue that ARC can be used to measure a human-like form of general fluid intelligence and that it enables fair general intelligence comparisons between AI systems and humans.

https://arxiv.org/abs/1911.01547

Now he seems to backtrack, with the release of harder ARC-like benchmarks, implying that the first one didn't actually test for really general human-like intelligence.

This sounds a bit like saying that a machine beating chess would require general intelligence -- but then adding, after Deep Blue beats chess, that chess doesn't actually count as a test for AGI, and that Go is the real AGI benchmark. And after a narrow system beats Go, moving the goalpost to beating Atari, and then to beating StarCraft II, then to MineCraft, etc.

At some point, intuitively real "AGI" will be necessary to beat one of these increasingly difficult benchmarks, but only because otherwise yet another benchmark would have been invented. Which makes these benchmarks mostly post hoc rationalizations.

A better approach would be to question what went wrong with coming up with the very first benchmark, and why a similar thing wouldn't occur with the second.

Re: François Chollet: The Arc Prize and How We Get to AGI [video]

#143
post #128
post #124

This may be a silly question, I'm no expert. But why not simply define as AGI any system that can answer a question that no human can. So for example, ask AGI to find out, from current knowledge, how to reconcile gravity and qed.

"What is the meaning of life, the universe, and everything?"

42

Re: François Chollet: The Arc Prize and How We Get to AGI [video]

#144
post #4

I feel like I'm the only one who isn't convinced getting a high score on the ARC eval test means we have AGI. It's mostly about pattern matching (and some of it ambiguous even for humans what the actual true response aught to be). It's like how in humans there's lots of different 'types' of intelligence, and just overfitting on IQ tests doesn't in my mind convince me a person is actually that smart.

In the video, François Chollet, creator of the ARC benchmarks, says that beating ARC does not equate to AGI. He specifically says they will be able to be beaten without AGI.

He only says this because otherwise he would have to say that

- OpenAI's o3 counts as "AGI" when it did unexpectedly beat the ARC-AGI benchmark or

- Explicitly admit that he was wrong when assuming that ARC-AGI would test for AGI

Re: François Chollet: The Arc Prize and How We Get to AGI [video]

#145
post #6
post #4

I feel like I'm the only one who isn't convinced getting a high score on the ARC eval test means we have AGI. It's mostly about pattern matching (and some of it ambiguous even for humans what the actual true response aught to be). It's like how in humans there's lots of different 'types' of intelligence, and just overfitting on IQ tests doesn't in my mind convince me a person is actually that smart.

I think the people behind the ARC Prize agree that getting a high score doesn't mean we have AGI. (They already updated the benchmark once to make it harder.) But an AGI should get a similarly high score as humans do. So current models that get very low scores are definitely not AGI, and likely quite far away from it.

> I think the people behind the ARC Prize agree that getting a high score doesn't mean we have AGI

The benchmark was literally called ARC-AGI. Only after OpenAI cracked it, they started backtracking and saying that it doesn't test for true AGI. Which undermines the whole premise of a benchmark.

Re: François Chollet: The Arc Prize and How We Get to AGI [video]

#146
post #124

This may be a silly question, I'm no expert. But why not simply define as AGI any system that can answer a question that no human can. So for example, ask AGI to find out, from current knowledge, how to reconcile gravity and qed.

Aside from other objections already mentioned, your example would require feasible experiments for verification, and likely the process of finding a successful theory of quantum gravity requires a back and forth between experimenters and theorists.

Re: François Chollet: The Arc Prize and How We Get to AGI [video]

#147
post #120

There is some kind of massive brigading happening on this thread. Lots of thoughtful comments are downmodded or flagged (including mine, which I thought was pretty thoughtful. I even said poop instead of shit.). https://news.ycombinator.com/item?id=44492241 My comment was basically instantly flagged. I see at least 3 other flagged comments that I can't imagine deserve to be flagged.

You didn’t address anything from the actual talk.

I addressed the entire concept of the talk, and made other relevant points. The correct response to "let me tell you something I can't possibly know" isn't to argue the points within that frame.

If you see a talk like: "How we will develop diplomacy with the rat-people of TRAPPIST-5." you don't have to make some argument about super-earths and gravity and the rocket equation. You can just point out it's absurd to pretend to know something like whether there are rat-people there.

Either way, it isn't flag-able!

Re: François Chollet: The Arc Prize and How We Get to AGI [video]

#148
post #120

Earlier quoted context omitted.

You didn’t address anything from the actual talk.

I addressed the entire concept of the talk, and made other relevant points. The correct response to "let me tell you something I can't possibly know" isn't to argue the points within that frame. If you see a talk like: "How we will develop diplomacy with the rat-people of TRAPPIST-5." you don't have to make some argument about super-earths and gravity and the rocket equation. You can just point out it's absurd to pre…

Did you actually watch the talk?

The flagging is probably due to your aggressively indignant style.

Re: François Chollet: The Arc Prize and How We Get to AGI [video]

#149

Earlier quoted context omitted.

In the video, François Chollet, creator of the ARC benchmarks, says that beating ARC does not equate to AGI. He specifically says they will be able to be beaten without AGI.

He only says this because otherwise he would have to say that - OpenAI's o3 counts as "AGI" when it did unexpectedly beat the ARC-AGI benchmark or - Explicitly admit that he was wrong when assuming that ARC-AGI would test for AGI

FWIW the original ARC was published in 2019, just after GPT-2 but a while before GPT-3. I work in the field, I think that discussing AGI seriously is actually kind of a recent thing (I'm not sure I ever heard the term 'AGI' until a few years ago). I'm not saying I know he didn't feel that, but he doesn't talk in such terms in the original paper.

Re: François Chollet: The Arc Prize and How We Get to AGI [video]

#150

Earlier quoted context omitted.

We don't really have a true test that means "if we pass this test we have AGI" but we have a variety of tests (like ARC) that we believe any true AGI would be able to pass. It's a "necessary but not sufficient" situation. Also ties directly to the challenge in defining what AGI really means. You see a lot of discussions of "moving the goal posts" around AGI, but as I see it we've never had goal posts, we've just got…

I don't think we actually even have a good definition of "This is what AGI is, and here are the stationary goal posts that, when these thresholds are met, then we will have AGI". If you judged human intelligence by our AI standards, then would humans even pass as Natural General Intelligence? Human intelligence tests are constantly changing, being invalidated, and rerolled as well. I maintain that today's modern LLMs…

>I don't think we actually even have a good definition of "This is what AGI is, and here are the stationary goal posts that, when these thresholds are met, then we will have AGI".

Not only do we not have that, I don't think it's possible to have it.

Philosophers have known about this problem for centuries. Wittgenstein recognized that most concepts don't have precise definitions but instead behave more like family resemblances. When we look at a family we recognize that they share physical characteristics, even if there's no single characteristic shared by all of them. They don't need to unanimously share hair color, skin complexion, mannerisms, etc. in order to have a family resemblance.

Outside of a few well-defined things in logic and mathematics, concepts operate in the same way. Intelligence isn't a well-defined concept, but that doesn't mean we can't talk about different types of human intelligence, non-human animal intelligence, or machine intelligence in terms of family resemblances.

Benchmarks are useful tools for assessing relative progress on well-defined tasks. But the decision of what counts as AGI will always come down to fuzzy comparisons and qualitative judgments.

Post reply on HN