Live data from Hacker News

Arc-AGI-2 and ARC Prize 2025

arcprize.org

71–80 of 103 posts

Re: Arc-AGI-2 and ARC Prize 2025

#71
post #2

Hey HN, Greg from ARC Prize Foundation here. Alongside Mike Knoop and François Francois Chollet, we’re launching ARC-AGI-2, a frontier AI benchmark that measures a model’s ability to generalize on tasks it hasn’t seen before, and the ARC Prize 2025 competition to beat it. In Dec ‘24, ARC-AGI-1 (2019) pinpointed the moment AI moved beyond pure memorization as seen by OpenAI's o3. ARC-AGI-2 targets test-time reasoning.…

> Our belief is that once we can no longer come up with quantifiable problems that are "feasible for humans and hard for AI" then we effectively have AGI. I don’t think that follows. Just because people fail to create ARC-AGI problems that are difficult for an AI to solve, doesn’t mean that said AI can just be plugged into a humanoid robot and it will now reliably cook dinner, order a pizza and drive to pick it up, t…

The task you mention require intelligence but also a robot body with a lot of physical dexterity suited to a designed-for-humanoids world. That seems like an additional requirement on top of intelligence? Maybe we do not want an AGI definition to include that?

There are humans who cannot perform these tasks, at least without assistive/adapted systems such as a wheelchair and accessible bus.

Re: Arc-AGI-2 and ARC Prize 2025

#72
post #2

Hey HN, Greg from ARC Prize Foundation here. Alongside Mike Knoop and François Francois Chollet, we’re launching ARC-AGI-2, a frontier AI benchmark that measures a model’s ability to generalize on tasks it hasn’t seen before, and the ARC Prize 2025 competition to beat it. In Dec ‘24, ARC-AGI-1 (2019) pinpointed the moment AI moved beyond pure memorization as seen by OpenAI's o3. ARC-AGI-2 targets test-time reasoning.…

Thanks for your awesome work Greg! The success of o3 directly contradicts us being in an "idea-constrained environment", what makes you believe that?

What makes you think so?

From ChatGPT 3.5 to o1, all LLMs progress came from investment in training: either by using much more data, or using higher quality data thanks to artificial data.

o1 (and then o3) broke this paradigm by applying a novel idea (RL+search on CoT) and that's because of it that it was able to make progress on ARC-AGI.

So IMO the success of o3 goes in favor of the argument of how we are in an idea-constrained environment.

Re: Arc-AGI-2 and ARC Prize 2025

#73
post #2

Hey HN, Greg from ARC Prize Foundation here. Alongside Mike Knoop and François Francois Chollet, we’re launching ARC-AGI-2, a frontier AI benchmark that measures a model’s ability to generalize on tasks it hasn’t seen before, and the ARC Prize 2025 competition to beat it. In Dec ‘24, ARC-AGI-1 (2019) pinpointed the moment AI moved beyond pure memorization as seen by OpenAI's o3. ARC-AGI-2 targets test-time reasoning.…

Using AGI in the titles of your tests might not be accurate or appropriate. May I suggest NAI - Narrow AI?

Re: Arc-AGI-2 and ARC Prize 2025

#75

Earlier quoted context omitted.

It's useful to know what current AI systems can achieve with unlimited test-time compute resources. Ultimately though, the "spirit of the challenge" is efficiency, which is why we're specifically looking for solutions that are at least within 1-2 order of magnitude of cost from being competitive with humans. The Kaggle leaderboard is very resource-constrained, and on the public leaderboard you need to use less than $…

Efficiency sounds like a hardware problem as much as a software problem. $10000 in compute is a moving target, today's GPUs are much much better than 10 years ago.

> $10000 in compute is a moving target

And it's also irrelevant in some fields. If you solve a "protein folding" problem that was a blocker for a pharma company, that 10k is peanuts now.

Same for coding. If you can spend 100$ / hr on a "mid-level" SWE agent but you can literally spawn 100 today and 0 tomorrow and reach your clients faster, again the cost is irrelevant.

Re: Arc-AGI-2 and ARC Prize 2025

#76
post #31
post #26

Earlier quoted context omitted.

If you aren't joking, that will filter most humans.

They said at least two people out of 400 solved each problem so they're pretty hard.

I don't think that's correct. They had 400 people receive some questions, and only kept the questions that were solved by at least 2 people. The 400 people didn't all receive 120 questions (they'd have probably got bored).

If you go through the example problems you'll notice that most are testing the "aha" moment. Once you do a couple, you know what to expect, but with larger grids you have to stay focused and keep track of a few things to get it right.

Re: Arc-AGI-2 and ARC Prize 2025

#77
post #73
post #2

Hey HN, Greg from ARC Prize Foundation here. Alongside Mike Knoop and François Francois Chollet, we’re launching ARC-AGI-2, a frontier AI benchmark that measures a model’s ability to generalize on tasks it hasn’t seen before, and the ARC Prize 2025 competition to beat it. In Dec ‘24, ARC-AGI-1 (2019) pinpointed the moment AI moved beyond pure memorization as seen by OpenAI's o3. ARC-AGI-2 targets test-time reasoning.…

Using AGI in the titles of your tests might not be accurate or appropriate. May I suggest NAI - Narrow AI?

My prediction: we'll be arguing about what AGI actually is... Forever.

Re: Arc-AGI-2 and ARC Prize 2025

#78

Earlier quoted context omitted.

They are useful to reach Arc-N+1

How are any of these a useful path to asking an AI to cook dinner? We already know many tasks that most humans can do relatively easily, yet most people don’t expect AI to be able to do them for years to come (for instance, L5 self-driving). ARC-AGI appears to be going in the opposite direction - can these models pass tests that are difficult for the average person to pass. These benchmarks are interesting in that th…

The "everyday tasks" you specifically mention involve motor skills that are not useful for measuring intelligence.

Re: Arc-AGI-2 and ARC Prize 2025

#79
post #73

Earlier quoted context omitted.

Using AGI in the titles of your tests might not be accurate or appropriate. May I suggest NAI - Narrow AI?

My prediction: we'll be arguing about what AGI actually is... Forever.

Or depending on your outlook, for a couple of years, and then we will no longer be participating in these or any other cognitive exercises.

Re: Arc-AGI-2 and ARC Prize 2025

#80
post #2

Hey HN, Greg from ARC Prize Foundation here. Alongside Mike Knoop and François Francois Chollet, we’re launching ARC-AGI-2, a frontier AI benchmark that measures a model’s ability to generalize on tasks it hasn’t seen before, and the ARC Prize 2025 competition to beat it. In Dec ‘24, ARC-AGI-1 (2019) pinpointed the moment AI moved beyond pure memorization as seen by OpenAI's o3. ARC-AGI-2 targets test-time reasoning.…

Thanks for your awesome work Greg! The success of o3 directly contradicts us being in an "idea-constrained environment", what makes you believe that?

Not Greg/team, so unrelated opinion. o3 solution for ARC v1 was incredibly expensive. Some good ideas are at least needed to take that cost down by a factor 100-10000x.
Post reply on HN