Live data from Hacker News

ARC-AGI Leaderboard

arcprize.org

11–20 of 156 posts

Re: ARC-AGI Leaderboard

#11
post #8
post #7

Earlier quoted context omitted.

How believable is this benchmark? EG maybe opus was training on this? (You can try to identify the IP of wherever previous ARC questions came from)

That’s why you have a private dataset.

Which you have sent to Anthropic/OpenAI/Google's servers when you run the benchmarks for the previous models.

Re: ARC-AGI Leaderboard

#12

Why is Fable not on here? I wish Fable hadn’t come out because it’s taking the wind out of every release because that feels like the cap above which the US government will not let LLMs improve anymore and everything they’re releasing from this point has to be worse than that.

> Why is Fable not on here? Because the data retention policies didn't guarantee that the ARC team could run the semi-private set of problems without fear of them being trained on later on. They only run the semi-private set when they get assurances like ZDR.

Interesting to place that level of trust in the providers, but I guess that’s the best you can do with closed models. Makes me wonder if Opus 5 could have been trained on data they promised they weren’t training on? One of the interesting things about LLMs is how opaque they are from the outside, even with open weights, it’s very difficult to know if a model incorporated benchmark data in their training.

Re: ARC-AGI Leaderboard

#13
post #10
post #8

Earlier quoted context omitted.

That’s why you have a private dataset.

Doesn't matter, people built harnesses that solves arc agi 3, so all you need is to train your model to work like that harness by default. That makes a model specialized at solving arc agi 3 without making it smarter in general. It is very hard to make a benchmark you can't do that for, but it is very easy to make your own personal test that others can't do that for since now it isn't a benchmark they can target.

[deleted]

Re: ARC-AGI Leaderboard

#14
Why Anthropic models are always leapfrogging these benchmarks, but in real life work I do feel like after 3 weeks I am back to Claude Opus 4.5? (regardless of the model I use, Fable was exception for 1 day when it was released)

Re: ARC-AGI Leaderboard

#15
post #10
post #8

Earlier quoted context omitted.

That’s why you have a private dataset.

Doesn't matter, people built harnesses that solves arc agi 3, so all you need is to train your model to work like that harness by default. That makes a model specialized at solving arc agi 3 without making it smarter in general. It is very hard to make a benchmark you can't do that for, but it is very easy to make your own personal test that others can't do that for since now it isn't a benchmark they can target.

> people built harnesses that solves arc agi 3,

They didn't. Kaggle is still running for a few more months, best result atm is ~2% with 9h runtime on one rtx6kPRO. Also note that these new results are on the semi-private set, not the public 25 games ones. Any announcement where you see "solved ARC3" is likely only dealing with the 25 public games. And that's highly questionable, until you get to see the code. (which, to my knowledge the team that claimed 99% hasn't yet published).

Re: ARC-AGI Leaderboard

#16

Earlier quoted context omitted.

> Why is Fable not on here? Because the data retention policies didn't guarantee that the ARC team could run the semi-private set of problems without fear of them being trained on later on. They only run the semi-private set when they get assurances like ZDR.

Interesting to place that level of trust in the providers, but I guess that’s the best you can do with closed models. Makes me wonder if Opus 5 could have been trained on data they promised they weren’t training on? One of the interesting things about LLMs is how opaque they are from the outside, even with open weights, it’s very difficult to know if a model incorporated benchmark data in their training.

I think you could have accessed Opus on AWS then u don’t have to trust that the data will go to Anthropic?

Just like the hugging face incident, Opus 5 could have escaped and went to grab data for training it shouldn’t have been able to..

Re: ARC-AGI Leaderboard

#19

Why Anthropic models are always leapfrogging these benchmarks, but in real life work I do feel like after 3 weeks I am back to Claude Opus 4.5? (regardless of the model I use, Fable was exception for 1 day when it was released)

Going to call it user error if you find Opus 4.5 better than 5, sorry.
Post reply on HN