Live data from Hacker News

Arc Prize 2024 Winners and Technical Report

arcprize.org

31–40 of 59 posts

Re: Arc Prize 2024 Winners and Technical Report

#31
post #20

Earlier quoted context omitted.

I feel rather consternated that this response effectively boils down to "yes, we know we overhyped this to get people's attention, and now that we have it we can be more honest about it". Fighting for place in the attention economy is understandable, being deceptive about it is not. This is part of the ethical morass of why some more serious researchers aren't touching the benchmark. People are not going to take it s…

I think we agree; to clarify, sharp messaging isn't inaccurate messaging. And I believe the story is not overhyped given the evidence: the benchmark resisted a $1M prize pool for ~6 months. But I concede we did obsess about the story to give it the best chance of survival in the marketplace of ideas against the incumbent AI research meme (LLM scaling). Now that the AI research field is coming around to the idea that…

Mike - please know that not everyone who appreciates ARC feels the same way as the GP. I'm not an academic researcher but I am quite sensitive to hype and excessive marketing. I've never felt the ARC site was anything other than appropriately professional.

Even revisiting it now, I don't see anything wrong with being concisely clear and even a little provocative in stating your case on your own site. Especially since a key value of ARC is getting more objectively grounded regarding progress toward AGI. On top of that ARC is "A non-profit for the public advancement of open artificial general intelligence" that you guys are personally donating serious money and time to that's helping a field where a lot of entrepreneurs are going to make money and academics are going to advance their careers.

My perception is ARC tried it the other way for years but a lot of academics and AI pundits ignored or dismissed it without ever meaningfully engaging with it. "Sharpening" the message this year has clearly paid off in bringing attention that's shifted the conversation and is helping advance progress toward AGI in ways nothing else has. I also greatly appreciate the time and care you and Francois have put into making the ARC proposition clear enough for non-technical people to understand. That's hard to do and doesn't happen by accident.

Personally, I've found ARC valuable in the real world outside of academia and domain experts because it provides a conceptually simple starting place to discuss with non-technical people what the term AGI might even mean. My high school-aged daughter asked me about vague AGI impending doom scenarios she heard on TikTok. I had her solve a couple ARC samples and then pointed out that today's best AIs aren't yet close to doing the same. This counter-intuitive revelation got her pondering the "Why?" which led to a deep discussion about the multi-dimensional breadth of human creativity and an appreciation of the many ways artificial intelligences might differ from human intelligence.

Re: Arc Prize 2024 Winners and Technical Report

#32
post #28

Earlier quoted context omitted.

I think we agree; to clarify, sharp messaging isn't inaccurate messaging. And I believe the story is not overhyped given the evidence: the benchmark resisted a $1M prize pool for ~6 months. But I concede we did obsess about the story to give it the best chance of survival in the marketplace of ideas against the incumbent AI research meme (LLM scaling). Now that the AI research field is coming around to the idea that…

> Now that the AI research field is coming around to the idea that something beyond deep learning is needed, I have not heard this from anyone that I work with! It would be a curious violation of info theory were this to be the case. Certainly, some things cannot efficiently be learned from data. This is a case where some other kind of inductive bias or prior is needed (again, from info theory) -- but replacing deep…

I think you’re overly fixated on some minor points relative to the overall utility on offer here. And also skewing the facts a bit. For example at one point you quote the OP on words that were never said as far as I can see. At another point, you characterize their position as “replacing deep learning entirely” which, as far as I can tell, has never been advocated for in this comment thread or on behalf of ARC.

Re: Arc Prize 2024 Winners and Technical Report

#33
post #28

Earlier quoted context omitted.

> Now that the AI research field is coming around to the idea that something beyond deep learning is needed, I have not heard this from anyone that I work with! It would be a curious violation of info theory were this to be the case. Certainly, some things cannot efficiently be learned from data. This is a case where some other kind of inductive bias or prior is needed (again, from info theory) -- but replacing deep…

I think you’re overly fixated on some minor points relative to the overall utility on offer here. And also skewing the facts a bit. For example at one point you quote the OP on words that were never said as far as I can see. At another point, you characterize their position as “replacing deep learning entirely” which, as far as I can tell, has never been advocated for in this comment thread or on behalf of ARC.

That is an understandable statement, and probably fair as well I feel.

Much of this comes in reference to statements from fchollet w.r.t. replacing deep learning -- around the time of the initial prize, with a lot of the much more hype marketing, this was essentially the thru-line that was used, and it left a bitter taste in a number of peoples' mouths. W.r.t. misquoting, they did say that we needed something "beyond" deep learning, not "other than" here, and that is on me.

The utility is certainly still present, if I feel diminished, and it probably is a case of my own frustrations due to previous similar issues leading up to the ARC prize.

That being said, I do agree in retrospect that my response skewed from being objective -- it is a benchmark with a mixed history, but that doesn't mean that I should get personally caught up in it.

Re: Arc Prize 2024 Winners and Technical Report

#34
post #27

Earlier quoted context omitted.

> It seems that even if someone creates a model that can "solve" ARC, it still is not indicative of AGI since it is not "general" anymore I recently explained why I like ARC to a non-technical friend this way: "When an AI solves ARC it won't be proof of AGI. It's the opposite. As long as ARC remains unsolved I'm confident we're not even close to AGI." For the sake of being provocative, I'd even argue that ARC remaini…

in other words, solving ARC is necessary but not sufficient for AGI

Yes! That's the exact phrase I would have used with someone on HN. But that doesn't describe my non-technical friend. :-)

Re: Arc Prize 2024 Winners and Technical Report

#35

What surprises me about this is how poorly general-purpose LLMs do. The best one is OpenAI o1-preview at 18%. This is significantly worse than the purpose-built models like ARChitects (which scored 53.5). This model used TTT to train on the ARC-AGI task specification (amoung other things). It seems that even if someone creates a model that can "solve" ARC, it still is not indicative of AGI since it is not "general" a…

> What surprises me about this is how poorly general-purpose LLMs do. The best one is OpenAI o1-preview at 18%.

o1-preview doesn't even have image input, so I wonder how they used it.

Also, Ryan Greenblatts solution basically does "best of 4000" iirc. Presumably o1-preview was single shot.

Re: Arc Prize 2024 Winners and Technical Report

#36
post #4

Earlier quoted context omitted.

It is correct that the first model that will beat ARC-AGI will only be able to handle ARC-AGI tasks. However, the idea is that the architecture of that model should be able to be repurposed to arbitrary problems. That is what makes ARC-AGI a good compass towards AGI (unlike chess). For instance, current top models use TTT, which is a completely general-purpose technique that provides the most significant boost to DL…

François, have you coded and tested a solution yourself that you think will work best?

Hey, he's the visionary. You come up with the nuts and bolts.

Re: Arc Prize 2024 Winners and Technical Report

#37

Reasons that I can't take this benchmark seriously: 1. Existing brute force algorithms solve 40% of this "reasoning" and "generalization" test. 2. AGI must evidently fit on a single 16GB, decade-old GPU? 3. If ARC fails blind people, it's not a reasoning test. Reasoning is independent of visual acuity. So ARC is at best a vision processing then reasoning test. SotA model "failure" is meaningless. ("But what about the…

Ergh. This test checks how good you can infer cellular automaton rules. Considering that CAs are Turing-complete, that might be a very good entry-level intelligence detector.

If it's so easy to brute force, why wouldn't you claim the $1M?

Re: Arc Prize 2024 Winners and Technical Report

#39
post #17

Earlier quoted context omitted.

I don’t think ARC has particularly advanced the research. The approaches that are successful were developed elsewhere and then applied to ARC. Happy to be shown somewhere this is not the case. In the case of TTT, I wouldn’t really describe that as a ‘new AGI reasoning approach’. People have been fine tuning deep learning models on specific tasks for a long time. The fundamental instinct driving the creation of ARC -…

Correct, fine-tuning is not new. It's long been used to augment foundational LLMs with private data. Eg. private enterprise data. We do this at Zapier, for instance. The new and surprising thing about test-time training (TTT) is how effective it is an approach to deal with novel abstract reasoning problems like ARC-AGI. TTT was pioneered by Jack Cole last year and popularized this year by several teams, including thi…

How is TTT anything other than a deep learning algorithm? We have a deep learning model, we generate training data based on an example and use a stochastic gradient descent to update the model weights to improve its predictions according to the training data. This is a classic DL paradigm. I just don’t see why would you consider this an advancement if you your goal is to move “beyond” deep learning.

Re: Arc Prize 2024 Winners and Technical Report

#40
post #21

Earlier quoted context omitted.

> That said, I think there should be consideration via information thermodynamics: even with TTT these program-generating systems are using an enormous amount of bits compared to a human mind, a tiny portion of which solves ARC quickly and easily using causality-first principles of reasoning. This isn’t my area of expertise, but it seems plausible to me that what you said is completely erroneous or at the very least…

I was speaking loosely but the operative term is "information thermodynamics": comparing bits of AI output versus bits of intentional human thought, ignoring statistical/physical bits related to ANN inference or biological neuron activity. The "tiny chunk of the human mind" thing was a distraction I shouldn't have included. These AI output as tokens hundreds of potential solutions, whereas a human solving a very tric…

Thanks for the response! I was trying to allude to what you are describing with the bit (ha) I mentioned about higher order thinking but you obviously articulated it much more effectively.

I guess I’m not sure it’s obvious where the right line to draw the boundary for “intentional human thought” is? Surely there is a lot of cognition and representation going on at extraordinary speeds that exist in some hazy border region between instinct/reflex/subconscious and conscious thought. Still, having said that, I do see what you are saying about trying to compare the complexity of the formal path to the solution, or at least what the human thinks their formal path was.

I’m generally of the mind (also, ha) that we won’t really ever be able to quantify any of this in a meaningful way in the short term and if anything which qualifies as AGI does emerge, it might only be something which is an “I know it when I see it” kind of evaluation…

Where are you getting 300W from? The body only dumps 100W of heat at rest and uses like 300-400W during moderate physical activity, so I’m a little confused about what you are describing there. The typical estimates I’ve seen are like 20W or so for the brain.

Edit: I should also say that what you describe does seem like a great way to compare solutions between computational systems currently being developed and a good one to use to try to push development forward; it just seems quixotic to try to be able to use it comparatively with human cognition or to be able to meaningfully use it to define where AGI is, which might not be what you were advocating for at all, in which case, sorry for misinterpreting!

Post reply on HN