Live data from Hacker News

ARC Prize – a $1M+ competition towards open AGI progress

arcprize.org

281–290 of 351 posts

Re: ARC Prize – a $1M+ competition towards open AGI progress

#281
post #268

Earlier quoted context omitted.

> The practice of solving problems that you describe is to ingrain/memorize those steps so you don't forget how to apply the procedure correctly Simply memorizing sequences of steps is not how mathematics learning works, otherwise we would not see so much variation in outcomes. Me and Terence Tao on the same exact math training data would not yield two mathematicians of similar skill. While it's true that memorizatio…

> Simply memorizing sequences of steps is not how mathematics learning works, otherwise we would not see so much variation in outcomes Everyone starts by memorizing how to do basic arithmetic on numbers, their multiplication tables and fractions. Only some then advance to understanding why those operations must work as they do. > It's worth noting that for composition, key to abstract reasoning, LLMs failed to genera…

I started by understanding. I could multiply by repeat addition (each addition counted one at a time with the aid of fingers) before I had the 10x10 addition table memorized. I learned university level calculus before I had more than half of the 10x10 multiplication table memorized, and even that was from daily use, not from deliberate memorization. There wasn't a day in my life where I could recite the full table.

Maybe schools teach by memorization, but my mom taught me by explaining what it means, and I highly recommend this approach (and am a proof by example that humans can learn this way).

Re: ARC Prize – a $1M+ competition towards open AGI progress

#282

Earlier quoted context omitted.

> why not just make the grid 1x1 and select a single color? For two reasons: 1. The initially suggested grid size was 3x3. 2. Filling in a 3x3 grid is sufficient to show that you understood the pattern, but filling in a 1x1 (or even 2x2) grid is insufficient. Requiring the user fill in a larger grid is a waste of time. The existence of the grid size selector would still make sense in cases where a 2x2 grid would be s…

The fact that two intelligent beings are debating what the correct answer is shows that there is no fixed correct answer that proves "intelligence". This is IQ tests all over again. Actually testing how alike you think to the author of the test.

Twist: bigyikes is an LLM!

Re: ARC Prize – a $1M+ competition towards open AGI progress

#283
post #196

Earlier quoted context omitted.

In another comment I replied that 3D high fidelity images do end up being thousands of training samples, so the answer is yes.

Are you suggesting that if a group of kids were given a book of zoo animals before going to the zoo, they would have difficulties identifing any new animals, because they only have seen one picture of each?

I think that's an interesting question, and a possible counter to my argument.

Certainly kids learn and become better at extrapolation and need fewer and fewer samples in general as they get more life experience.

Re: ARC Prize – a $1M+ competition towards open AGI progress

#284

> the only eval which measures AGI. That's a stretch. This is a problem at which LLMs are bad. That does not imply it's a good measure of artificial general intelligence. After working a few of the problems, I was wondering how many different transformation rules the problem generator has. Not very many, it seems. So the problem breaks down into extracting the set of transformation rules from the data, then applying…

Yes to your last question, that is essentially how the first iteration solutions operated. Some of the original kaggle competition’s best solutions used a DSL made of these transformations. That was 4 years ago. [1]

The issue with that path is that the problems aren’t using a programmatic generator. The rule sets are anything a person could come up with. It might be as simple as “biggest object turns blue” but they can be much more complicated.

Additionally, the test set is private so it can’t be trained on or extracted from. It has rules that aren’t in the public sets.

[1] https://www.kaggle.com/competitions/abstraction-and-reasonin...

Re: ARC Prize – a $1M+ competition towards open AGI progress

#285
post #195

Earlier quoted context omitted.

Are you saying it's not fair for LLMs, because of the way they are taught is different? The difference is that we don't know better methods for them, but we do know of better methods for people.

I think they're saying that it's silly to claim humans learn with less data than LLMs, when humans are ingesting a continuous video, audio, olfactory and tactile data stream for 16+ hours a day, every day. It takes at least 4 years for a human children to be in any way comparable in performance to GPT-4 on any task both of them could be tested on; do people really believe GPT-4 was trained with more data than a 4 yea…

> I think they're saying that it's silly to claim humans learn with less data than LLMs, when humans are ingesting a continuous video, audio, olfactory and tactile data stream for 16+ hours a day, every day.

Yeah, but they're seeing mostly the same thing day after day!

They aren't seeing 10k stills of 10k different dogs, then 10k stills of 10k different cats. They're seeing $FOO thousand images of the family dog and the family cat.

My (now 4.5yo) toddler did reliably tell the difference between cats and dogs the first time he went with us to the local SPCA and saw cats and dogs that were not our cats and dogs.

In effect, 2 cats and 2 dogs were all he needed to reliably distinguish between cats and dogs.

Re: ARC Prize – a $1M+ competition towards open AGI progress

#286

Earlier quoted context omitted.

> Now, is it 10k examples? No, but I think it was on the order of hundreds, if not thousands. I have kids so I'm presuming I'm allowed to have an opinion here. This is ignoring the fact that babies are not just learning labels, they're learning the whole of language, motion planning, sensory processing, etc. Once they have the basics down concept acquisition time shrinks rapidly and kids can easily learn their new fa…

> kids can easily learn their new favorite animal in as little as a single example Until they encounter a similar animal and get confused, at which point you understand the implicit heuristic they were relying on. (Eg. They confused a dairy cow as a zebra, which means their heuristic was a black-and-white quadrupedal) Doesn't this seem remarkably close to how LLMs behave with one-shot or few-shot learning? I think th…

My favourite part of this is when they apply their new words to things that technically make sense, but don't. My daughter proudly pointed at a king wearing a crown as "sharp king" after learning about knives, saws, etc.

Re: ARC Prize – a $1M+ competition towards open AGI progress

#287
post #12

While I agree with the spirit of the competition, a $1M prize seems a little too low considering tens of billions of dollars have already been invested in the race to AGI, and we will see many times that put into the space in the coming years. The impact of AGI will be measured in trillions at minimum. So what you are ultimately rewarding isn't AGI research but fine tuning the newest public LLM release to best meet t…

The leaderboard is on the website. What medium should they use? https://arcprize.org/leaderboard

Re: ARC Prize – a $1M+ competition towards open AGI progress

#288

Earlier quoted context omitted.

> Simply memorizing sequences of steps is not how mathematics learning works, otherwise we would not see so much variation in outcomes Everyone starts by memorizing how to do basic arithmetic on numbers, their multiplication tables and fractions. Only some then advance to understanding why those operations must work as they do. > It's worth noting that for composition, key to abstract reasoning, LLMs failed to genera…

I started by understanding. I could multiply by repeat addition (each addition counted one at a time with the aid of fingers) before I had the 10x10 addition table memorized. I learned university level calculus before I had more than half of the 10x10 multiplication table memorized, and even that was from daily use, not from deliberate memorization. There wasn't a day in my life where I could recite the full table. M…

> I started by understanding. I could multiply by repeat addition

How did you learn what the symbols for numbers mean and how addition works? Did you literally just see "1 + 3 = 4" one day and intuit the meaning of all of those symbols? Was it entirely obvious to you from the get-go that "addition" was the same as counting using your fingers which was also the same as counting apples which was also the same as these little squiggles on paper?

There's no escaping the fact that there's memorization happening at some level because that's the only way to establish a common language.

Re: ARC Prize – a $1M+ competition towards open AGI progress

#289

Earlier quoted context omitted.

> Simply memorizing sequences of steps is not how mathematics learning works, otherwise we would not see so much variation in outcomes Everyone starts by memorizing how to do basic arithmetic on numbers, their multiplication tables and fractions. Only some then advance to understanding why those operations must work as they do. > It's worth noting that for composition, key to abstract reasoning, LLMs failed to genera…

The point is the memorization exercise requires orders of magnitude fewer examples for bootstrapping.

Does it though? It's a common claim but I don't think that's been rigourously established.

Re: ARC Prize – a $1M+ competition towards open AGI progress

#290

Earlier quoted context omitted.

I think they're saying that it's silly to claim humans learn with less data than LLMs, when humans are ingesting a continuous video, audio, olfactory and tactile data stream for 16+ hours a day, every day. It takes at least 4 years for a human children to be in any way comparable in performance to GPT-4 on any task both of them could be tested on; do people really believe GPT-4 was trained with more data than a 4 yea…

> I think they're saying that it's silly to claim humans learn with less data than LLMs, when humans are ingesting a continuous video, audio, olfactory and tactile data stream for 16+ hours a day, every day. Yeah, but they're seeing mostly the same thing day after day! They aren't seeing 10k stills of 10k different dogs, then 10k stills of 10k different cats. They're seeing $FOO thousand images of the family dog and…

> In effect, 2 cats and 2 dogs were all he needed to reliably distinguish between cats and dogs.

I assume he was also exposed to many images, photos and videos (realistic or animated) of cats and dogs in children books and toys he handled. In our case, this was a significant source of animal recognition skills of my daughters.

Post reply on HN