Live data from Hacker News

François Chollet: The Arc Prize and How We Get to AGI [video]

youtube.com

51–60 of 230 posts

Re: François Chollet: The Arc Prize and How We Get to AGI [video]

#51
post #19

Earlier quoted context omitted.

Today’s llms are fancy autocomplete but lack test time self learning or persistent drive. By contrast, an AGI would require: – A goal-generation mechanism (G) that can propose objectives without external prompts – A utility function (U) and policy π(a│s) enabling action selection and hierarchy formation over extended horizons – Stateful memory (M) + feedback integration to evaluate outcomes, revise plans, and execute…

I'd say we're not far off. Looking at the human side, it takes a while to actually learn something. If you've recently read something it remains in your "context window". You need to dream about it, to think about, to revisit and repeat until you actually learn it and "update your internal model". We need a mechanism for continuous weight updating. Goal-generation is pretty much covered by your body constantly drip-f…

Yes, you're right, that's what we're doing.

https://github.com/dmf-archive/PILF

Re: François Chollet: The Arc Prize and How We Get to AGI [video]

#54
post #4

I feel like I'm the only one who isn't convinced getting a high score on the ARC eval test means we have AGI. It's mostly about pattern matching (and some of it ambiguous even for humans what the actual true response aught to be). It's like how in humans there's lots of different 'types' of intelligence, and just overfitting on IQ tests doesn't in my mind convince me a person is actually that smart.

> It's mostly about pattern matching...

For all we know, human intelligence is just an emergent property of really good pattern matching.

Re: François Chollet: The Arc Prize and How We Get to AGI [video]

#55
post #24
post #4

I feel like I'm the only one who isn't convinced getting a high score on the ARC eval test means we have AGI. It's mostly about pattern matching (and some of it ambiguous even for humans what the actual true response aught to be). It's like how in humans there's lots of different 'types' of intelligence, and just overfitting on IQ tests doesn't in my mind convince me a person is actually that smart.

You're not alone in this; I expect us to have not yet enumerated all the things that we ourselves mean by "intelligence". But conversely, not passing this test is a proof of not being as general as a human's intelligence.

I find the "what is intelligence?" discussion a little pointless if I'm honest. It's similar to asking a question like does it mean to be a "good person" and would we know whether an AI or person is really "good"?

While understanding why a person or AI is doing what it's doing can be important (perhaps specifically in safety contexts) at the end of the day all that's really going to matter to most people is the outcomes.

So if an AI can use what appears to be intelligence to solve general problems and can act in ways that are broadly good for society, whether or not it meets some philosophical definition of "intelligent" or "good" doesn't matter much – at least in most contexts.

That said, my own opinion on this is that the truth is likely in between. LLMs today seem extremely good at being glorified auto-completes, and I suspect most (95%+) of what they do is just recalling patterns in their weights. But unlike traditional auto-completes they do seem to have some ability to reason and solve truly novel problems. As it stands I'd argue that ability is fairly poor, but this might only represent 1-2% of what we use intelligence for.

If I were to guess why this is I suspect it's not that LLM architecture today is completely wrong, but that the way LLMs are trained means that in general knowledge recall is rewarded more than reasoning. This is similar to the trade-off we humans have with education – do you prioritise the acquisition of knowledge or critical thinking? Maybe believe critical thinking is more important and should be prioritised more, but I suspect for the vast majority of tasks we're interested in solving knowledge storage and recall is actually more important.

Re: François Chollet: The Arc Prize and How We Get to AGI [video]

#56
post #41

Earlier quoted context omitted.

I've said this somewhere else, but we have the perfect test for AGI in the form of any open world game. Give the instructions to the AGI that it should finish the game and how to control it. Give the frames as input and wait. When I think of the latest Zelda games and especially how the Shrine chanllenges are desgined they especially feel like the perfect environement for an AGI test.

And if someone makes a machine that does all that and another person says "That's not really AGI because xyz" What then? The difficulty in coming up with a test for AGI is coming up with something that people will accept a passing grade as AGI. In many respects I feel like all of the claims that models don't really understand or have internal representation or whatever tend to lean on nebulous or circular definitions…

Given your premise (which I agree with) I think the issue in general comes from the lack of a good, broadly accepted definition of what AGI is. My initial comment originates from the fact that in my internal definition, an AGI would have a de facto understanding of the physics of "our world". Or better, could infer them by trial and error. But, indeed, it doesn't have to be the case. (The other advantage of the Zelda games is that they introduce new abilities that don't exist in our world, and for which most children -I've seen- understand the mechanisms and how they could be applied to solve a problem quite naturaly even they've never had that ability before).

Re: François Chollet: The Arc Prize and How We Get to AGI [video]

#57
post #4

I feel like I'm the only one who isn't convinced getting a high score on the ARC eval test means we have AGI. It's mostly about pattern matching (and some of it ambiguous even for humans what the actual true response aught to be). It's like how in humans there's lots of different 'types' of intelligence, and just overfitting on IQ tests doesn't in my mind convince me a person is actually that smart.

The point is not that having a high score -> AGI, their ideas are more of having a low score -> we don't have AGI yet.

Re: François Chollet: The Arc Prize and How We Get to AGI [video]

#58
post #4

I feel like I'm the only one who isn't convinced getting a high score on the ARC eval test means we have AGI. It's mostly about pattern matching (and some of it ambiguous even for humans what the actual true response aught to be). It's like how in humans there's lots of different 'types' of intelligence, and just overfitting on IQ tests doesn't in my mind convince me a person is actually that smart.

I understand Chollet is transparent that the "branding" of the ARC-AGI-n suites is meant to be suggestive of its purpose, than substantial. However, it does rub me the wrong way - as someone who's cynical of how branding can enable breathless AI hype by bad journalism. A hypothetical comparison would be labelling SHRDLU's (1968) performance on Block World planning tasks as "ARC-AGI-(-1)".[0] A less loaded name like (…

I get what you're saying about perception being reality and that ARC-AGI suggests beating it means AGI has been achieved.

In practice when I have seen ARC brought up, it has more nuance than any of the other benchmarks.

Unlike, Humanity's Last Exam, which is the most egregious example I have seen in naming and when it is referenced in terms of an LLMs capability.

Re: François Chollet: The Arc Prize and How We Get to AGI [video]

#59

The Arc prize/benchmark is a terrible judge of whether we got to AGI. If we assume that humans have "general intelligence", we would assume all humans could ace Arc... but they can't. Try asking your average person, i.e. supermarket workers, gas station attendants etc to do the Arc puzzles, they will do poorly, especially on the newer ones, but AI has to do perfectly to prove they have general intelligence? (not tryi…

Maybe it is a cultural difference aspect, but I feel that "supermarket workers, gas station attendants" (in an Asian country) that I know of should be quite capable of most ARC tasks.

Re: François Chollet: The Arc Prize and How We Get to AGI [video]

#60
post #33

Is the text available for those who don't hear so well?

At the very least, YouTube provides a transcript and a "Show Transcript" button in the video description, which you can click on to follow along.

When I watched the video I had the subtitles on. The automatic transcript is pretty good. "Test-time" which is used frequently gets translated as "Tesla" so watch out for that.
Post reply on HN