Note that this uses a harness so it doesn't qualify for the official ARC-AGI-3 leaderboard According to the authors the harness isn't ARC-AGI specific though https://x.com/agenticasdk/status/2037335806264971461
We're calling agents harnesses now?
Day 1 of ARC-AGI-3
31–40 of 79 posts
Re: Day 1 of ARC-AGI-3
#32Note that this uses a harness so it doesn't qualify for the official ARC-AGI-3 leaderboard According to the authors the harness isn't ARC-AGI specific though https://x.com/agenticasdk/status/2037335806264971461
It is 100% ARC-AGI-3 specific though, just read through the prompts https://github.com/symbolica-ai/ARC-AGI-3-Agents/blob/symbol...
Re: Day 1 of ARC-AGI-3
#33Note that this uses a harness so it doesn't qualify for the official ARC-AGI-3 leaderboard According to the authors the harness isn't ARC-AGI specific though https://x.com/agenticasdk/status/2037335806264971461
Doesn't the chat version of chatgpt or gemini also have interleaved tool calls, so do those also count as with harnesses?
Re: Day 1 of ARC-AGI-3
#34Re: Day 1 of ARC-AGI-3
#35Note that this uses a harness so it doesn't qualify for the official ARC-AGI-3 leaderboard According to the authors the harness isn't ARC-AGI specific though https://x.com/agenticasdk/status/2037335806264971461
> this uses a harness This seems like an arbitrary restriction. Tool-use requires a harness, and their whitepaper never defines exactly what counts as valid.
Re: Day 1 of ARC-AGI-3
#36Earlier quoted context omitted.
All traffic is monitored, all signal sources are eventually incorporated into the training set in one way or another. The person you're responding to is correct, even a single API call to any AI provider is sufficient to discount future results from the same provider.
You live in a conspiracy world. Those AI providers don't update the models that fast. You can try ask them solve ARC-AGI-3 without harness and see them struggle as yesterday yourself.
Re: Day 1 of ARC-AGI-3
#37Note that this uses a harness so it doesn't qualify for the official ARC-AGI-3 leaderboard According to the authors the harness isn't ARC-AGI specific though https://x.com/agenticasdk/status/2037335806264971461
It is 100% ARC-AGI-3 specific though, just read through the prompts https://github.com/symbolica-ai/ARC-AGI-3-Agents/blob/symbol...
Re: Day 1 of ARC-AGI-3
#38https://en.wikipedia.org/wiki/Goodhart's_law > Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.
Re: Day 1 of ARC-AGI-3
#39What if you give opus the same harness? Do people even care about meaningful comparisons any more or is it all just “numbers go up”
Re: Day 1 of ARC-AGI-3
#40we constantly underestimate the power of inference scaffolding. I have seen it in all domains: coding, ASR, ARC-AGI benchmarks you name it. Scaffolding can do a lot! And post-training too. I am confident our currently pre-trained models can beat this benchmark over 80% with the right post-training and scaffolding. That being said I don't think ARC-AGI proves much. It is not a useful task at all in the wild. it is jus…