Day 1 of ARC-AGI-3
11–20 of 79 posts
Re: Day 1 of ARC-AGI-3
#12On the public set of 25 problems. These are intended for development and testing, not evaluation. There are 110 private problems for actual evaluation purposes, and the ARC-AGI-3 paper says "the public set is materially easier than the private set".
Re: Day 1 of ARC-AGI-3
#13Note that this uses a harness so it doesn't qualify for the official ARC-AGI-3 leaderboard According to the authors the harness isn't ARC-AGI specific though https://x.com/agenticasdk/status/2037335806264971461
I for one think that harness development is perhaps the most interesting part at the moment and would love to have an alternative leaderboard with harnesses.
Re: Day 1 of ARC-AGI-3
#14On the public set of 25 problems. These are intended for development and testing, not evaluation. There are 110 private problems for actual evaluation purposes, and the ARC-AGI-3 paper says "the public set is materially easier than the private set".
Benchmarks on public tests are too easy to game. The model owners can just incorporate the answers in to the dataset. Only the private problems actually matter.
Re: Day 1 of ARC-AGI-3
#15Earlier quoted context omitted.
Benchmarks on public tests are too easy to game. The model owners can just incorporate the answers in to the dataset. Only the private problems actually matter.
In this case the code is public and you can see they are not cheating in that sense.
Re: Day 1 of ARC-AGI-3
#16Earlier quoted context omitted.
In this case the code is public and you can see they are not cheating in that sense.
Once the model has seen the questions and answers in the training stage, the questions are worthless. Only a test using previously unseen questions has merit.
Re: Day 1 of ARC-AGI-3
#17Earlier quoted context omitted.
Once the model has seen the questions and answers in the training stage, the questions are worthless. Only a test using previously unseen questions has merit.
They aren't training new models for this. This is an agent harness for Opus 4.6.
Re: Day 1 of ARC-AGI-3
#18Earlier quoted context omitted.
Benchmarks on public tests are too easy to game. The model owners can just incorporate the answers in to the dataset. Only the private problems actually matter.
In this case the code is public and you can see they are not cheating in that sense.
Re: Day 1 of ARC-AGI-3
#19Earlier quoted context omitted.
They aren't training new models for this. This is an agent harness for Opus 4.6.
All traffic is monitored, all signal sources are eventually incorporated into the training set in one way or another. The person you're responding to is correct, even a single API call to any AI provider is sufficient to discount future results from the same provider.
Re: Day 1 of ARC-AGI-3
#20Note that this uses a harness so it doesn't qualify for the official ARC-AGI-3 leaderboard According to the authors the harness isn't ARC-AGI specific though https://x.com/agenticasdk/status/2037335806264971461