Live data from Hacker News

ARC Prize – a $1M+ competition towards open AGI progress

arcprize.org

291–300 of 351 posts

Re: ARC Prize – a $1M+ competition towards open AGI progress

#291
post #266

Earlier quoted context omitted.

> The practice of solving problems that you describe is to ingrain/memorize those steps so you don't forget how to apply the procedure correctly Perhaps that is how you learned math, but it is nothing like how I learned math. Memorizing steps does not help, I sucked at it. What works for me us understanding the steps and why we used them. Once I understood the process and why it worked, I was able to reason my way th…

> Perhaps that is how you learned math, but it is nothing like how I learned math. Memorization is literally how you learned arithmetic, multiplication tables and fractions. Everyone starts learning math by memorization, and only later start understanding why certain steps work. Some people don't advance to that point, and those that do become more adept at math.

> Memorization is literally how you learned arithmetic, multiplication tables and fractions

I understood how to do arithmetic for numbers with multiple digits before I was taught a "procedure". Also, I am not even sure what you mean by "memorization is how you learned fractions". What is there to memorize?

Re: ARC Prize – a $1M+ competition towards open AGI progress

#292
post #126

I'm Simon Strandgaard and I participated in ARCathon 2022 (solved 3 tasks) and ARCathon 2023 (solved 8 tasks). I'm collecting data for how humans are solving ARC tasks, and so far collected 4100 interaction histories ( https://github.com/neoneye/ARC-Interactive-History-Dataset ). Besides ARC-AGI, there are other ARC like datasets, these can be tried in my editor ( https://neoneye.github.io/arc/ ). I have made some vi…

"Here is a challenge, designed to be unsolvable or so. We'll give you a bazillion dollars if you complete the challenge, and, in the meantime, we will use your attempts to train an as AI that will be worth the cost!!"

In the most charitable interpretation of this comment - I can understand the feeling, when so much of social media interactions are in the form 'It's post a picture of you as a baby, 10 year old, and current age!'. Those and many other instances can bring out excessive skepticism

But the people involved in this haven't signaled that they are in that path, either in the message about the challenge (precisely the opposite) or seemingly in their careers so far

So I guess I don't share the concern but a better way to phrase your comment could be -

"how can we be sure the human-provided solutions won't turn out to be just fodder for training a RL model or something that will later be monetized, closed and proprietary? Do the challenge organizers provide any guarantees on that?"

Re: ARC Prize – a $1M+ competition towards open AGI progress

#293

Earlier quoted context omitted.

The A in AGI stands for artificial, so a human+LLM system would not qualify as it has a natural, human component. That doesn't mean it's not an interesting topic, or that it won't help humans discover our world better, it's just the wrong label. Remove the human and you'd just have LLMs talking nonsense at each other. It's not surprising that you get an intelligent system when you include natural intelligence.

The key ingredients are not the humans but the feedback they carry to the model. Humans are embodied and can test ideas in the real world, LLMs need some kind of special deployment to achieve that. It just so happens that chat rooms are such a deployment. For example, AlphaZero started from scratch and only had feedback from the self-play game outcomes, but that was enough to reach superhuman level. It was the feedba…

That _is_ an interesting paper, I'll need to give it a read through.

Re: ARC Prize – a $1M+ competition towards open AGI progress

#294
post #108

Is there a leaderboard for the no-restriction version of the competition? I want to see how gpt4 does on it.

Yes there is a secondary leaderboard called ARC-AGI-Pub (in beta) with no limitations: https://arcprize.org/leaderboard

I don’t see gpt4 scores there. In fact I’m particularly interested in the performance of a natively multimodal model, like gpt4o or gemini. It does not really make sense to test a model trained on text on those visual/spatial puzzles.

Re: ARC Prize – a $1M+ competition towards open AGI progress

#295

Earlier quoted context omitted.

I haven't seen 1000 cats in my entire life. I'm sure I learned how to tell a dog from a cat after being exposed to just a single instance of each.

I'm sure you saw over 1B images of cats though, assuming 24 images per second from vision.

> I'm sure you saw over 1B images of cats though, assuming 24 images per second from vision.

The AI models aren't seeing the same image 1B times.

Re: ARC Prize – a $1M+ competition towards open AGI progress

#296

Earlier quoted context omitted.

> I think they're saying that it's silly to claim humans learn with less data than LLMs, when humans are ingesting a continuous video, audio, olfactory and tactile data stream for 16+ hours a day, every day. Yeah, but they're seeing mostly the same thing day after day! They aren't seeing 10k stills of 10k different dogs, then 10k stills of 10k different cats. They're seeing $FOO thousand images of the family dog and…

> In effect, 2 cats and 2 dogs were all he needed to reliably distinguish between cats and dogs. I assume he was also exposed to many images, photos and videos (realistic or animated) of cats and dogs in children books and toys he handled. In our case, this was a significant source of animal recognition skills of my daughters.

> I assume he was also exposed to many images, photos and videos (realistic or animated) of cats and dogs in children books and toys he handled.

No images or photos (no books).

TV, certainly, but I consider it unlikely that animals in the animation style of pepper pig helps the classifier.

Besides which, we're still talking under a dozen cats/dogs seen till that point.

Forget about cats/dogs. Here's another example: he only had to see a burger patty once to determine that it was an altogether new type of food, different from (for example) a sausage.

Anyone who has kids will have dozens of examples where the classifier worked without a false positive off a single novel item.

Re: ARC Prize – a $1M+ competition towards open AGI progress

#297

Earlier quoted context omitted.

https://arcprize.org/play?task=00576224 Yes the same puzzle. And I followed the second example. This was my solution: GRG OBO RGR B is the cyan like blue color. My solution looks right, but it says it’s wrong.

You need to resize the output grid to 6x6.

Yeah into the same problem but managed to find the solution by resizing it.

Re: ARC Prize – a $1M+ competition towards open AGI progress

#298

Earlier quoted context omitted.

> Perhaps that is how you learned math, but it is nothing like how I learned math. Memorization is literally how you learned arithmetic, multiplication tables and fractions. Everyone starts learning math by memorization, and only later start understanding why certain steps work. Some people don't advance to that point, and those that do become more adept at math.

> Memorization is literally how you learned arithmetic, multiplication tables and fractions I understood how to do arithmetic for numbers with multiple digits before I was taught a "procedure". Also, I am not even sure what you mean by "memorization is how you learned fractions". What is there to memorize?

> I understood how to do arithmetic for numbers with multiple digits before I was taught a "procedure"

What did you understand, exactly? You understood how to "count" using "numbers" that you also memorized? You intuitively understood that addition was counting up and subtraction was counting down, or did you memorize those words and what they meant in reference to counting?

> Also, I am not even sure what you mean by "memorization is how you learned fractions". What is there to memorize?

The procedure to add or subtract fractions by establishing a common denominator, for instance. The procedure for how numerators and denominators are multiplied or divided. I could go on.

Re: ARC Prize – a $1M+ competition towards open AGI progress

#299
post #103

Earlier quoted context omitted.

> Now, is it 10k examples? No, but I think it was on the order of hundreds, if not thousands. I have kids so I'm presuming I'm allowed to have an opinion here. This is ignoring the fact that babies are not just learning labels, they're learning the whole of language, motion planning, sensory processing, etc. Once they have the basics down concept acquisition time shrinks rapidly and kids can easily learn their new fa…

> How many homework questions did your entire calc 1 class have? I'm guessing less than 100… I’m quite surprised at this guess and intrigued by your school’s methodology. I would have estimated >30 problems average across 20 weeks for myself. My kids are still in pre-algebra, but they get way more drilling still, well over 1000 problems per semester once Zern, IReady, etc. are factored in. I believe it’s too much, bu…

I preferred doing large problem sets in math class because that is the only way I felt like I could gain an innate understanding of the math.

For example after doing several hundred logarithms, I was eventually able to do logs to 2 decimal places in my head. (Sadly I cannot do that anymore!) I imagine if I had just done a dozen or so problems I would not have gained that ability.

Re: ARC Prize – a $1M+ competition towards open AGI progress

#300

Chollet's argument is that LLMs just imitate and recombine patterns. This might be true if you're looking at LLMs in isolation, but when they chat with people something different happens. The system made of humans+LLMs is an AGI. It is no longer just a parrot, it ingests new information, gets guidance, feedback and is basically embodied in a chat room with human and tools. This scales for 200M users and 1 billion ses…

I fully, comprehensively agree with your take and have repeatedly arrived at the same conclusions in my research.
Post reply on HN