Live data from Hacker News

ARC-AGI Leaderboard

arcprize.org

21–30 of 156 posts

Re: ARC-AGI Leaderboard

#21

Why is Fable not on here? I wish Fable hadn’t come out because it’s taking the wind out of every release because that feels like the cap above which the US government will not let LLMs improve anymore and everything they’re releasing from this point has to be worse than that.

> Why is Fable not on here? Because the data retention policies didn't guarantee that the ARC team could run the semi-private set of problems without fear of them being trained on later on. They only run the semi-private set when they get assurances like ZDR.

How do they handle these assurances? Personally I have zero trust in the AI companies not trying to use this data to get ahead in the game, and short of sharing the weights and harness so that the benchmarkers can run the models themselves, I don't see a satisfactory solution with this mindset.

Re: ARC-AGI Leaderboard

#23
post #10

Earlier quoted context omitted.

Doesn't matter, people built harnesses that solves arc agi 3, so all you need is to train your model to work like that harness by default. That makes a model specialized at solving arc agi 3 without making it smarter in general. It is very hard to make a benchmark you can't do that for, but it is very easy to make your own personal test that others can't do that for since now it isn't a benchmark they can target.

> people built harnesses that solves arc agi 3, They didn't. Kaggle is still running for a few more months, best result atm is ~2% with 9h runtime on one rtx6kPRO. Also note that these new results are on the semi-private set, not the public 25 games ones. Any announcement where you see "solved ARC3" is likely only dealing with the 25 public games. And that's highly questionable, until you get to see the code. (which,…

For frontier models, not local.

https://schema-harness.github.io/

Re: ARC-AGI Leaderboard

#24

I have a suspicion that they are just trained on puzzles by now

There are private datasets, and 3rd party providers of these models. Fable doesn’t have a datapoint here because of its particular data retention policy. Even if you don’t trust AWS, do you think Opus on AWS is also sending the data to Anthropic? Do you have any evidence?

Re: ARC-AGI Leaderboard

#25

Why Anthropic models are always leapfrogging these benchmarks, but in real life work I do feel like after 3 weeks I am back to Claude Opus 4.5? (regardless of the model I use, Fable was exception for 1 day when it was released)

Well, what kinds of things do you see Opus 4.5 completely fail at? Maybe those are not the ones that newer models have improved on.

Re: ARC-AGI Leaderboard

#26

Why Anthropic models are always leapfrogging these benchmarks, but in real life work I do feel like after 3 weeks I am back to Claude Opus 4.5? (regardless of the model I use, Fable was exception for 1 day when it was released)

Enshittification.

Re: ARC-AGI Leaderboard

#27

Earlier quoted context omitted.

> people built harnesses that solves arc agi 3, They didn't. Kaggle is still running for a few more months, best result atm is ~2% with 9h runtime on one rtx6kPRO. Also note that these new results are on the semi-private set, not the public 25 games ones. Any announcement where you see "solved ARC3" is likely only dealing with the 25 public games. And that's highly questionable, until you get to see the code. (which,…

For frontier models, not local. https://schema-harness.github.io/

Yes, saw that. They haven't yet released any code. Until they do, treat it with a huuuge grain of salt. In fact treat any 99% result in ML with a huge grain of salt.

Re: ARC-AGI Leaderboard

#28

Why Anthropic models are always leapfrogging these benchmarks, but in real life work I do feel like after 3 weeks I am back to Claude Opus 4.5? (regardless of the model I use, Fable was exception for 1 day when it was released)

It's called frog boiling.

We get used to the new level of intelligence so fast, any deviation feels like going back to the stone age.

If you don't believe me, create something complex with Opus 5 and then with Opus 4.5, and notice the difference.

Re: ARC-AGI Leaderboard

#29

I have a suspicion that they are just trained on puzzles by now

There are private datasets, and 3rd party providers of these models. Fable doesn’t have a datapoint here because of its particular data retention policy. Even if you don’t trust AWS, do you think Opus on AWS is also sending the data to Anthropic? Do you have any evidence?

The fact that it says so in the licensing conditions on AWS?
Post reply on HN