Live data from Hacker News

Day 1 of ARC-AGI-3

symbolica.ai

31–40 of 79 posts

Re: Day 1 of ARC-AGI-3

#31
post #2

Note that this uses a harness so it doesn't qualify for the official ARC-AGI-3 leaderboard According to the authors the harness isn't ARC-AGI specific though https://x.com/agenticasdk/status/2037335806264971461

We're calling agents harnesses now?

The point of this test is to check if an AI system can figure out the game. This isn't what happened here. A human figured out the game, wrote in their prompts exactly how the game works and THEN put the AI on the problem. This is 100% cheating and imo quite stupid.

Re: Day 1 of ARC-AGI-3

#32
post #2

Note that this uses a harness so it doesn't qualify for the official ARC-AGI-3 leaderboard According to the authors the harness isn't ARC-AGI specific though https://x.com/agenticasdk/status/2037335806264971461

It is 100% ARC-AGI-3 specific though, just read through the prompts https://github.com/symbolica-ai/ARC-AGI-3-Agents/blob/symbol...

What a dick move. Making that prompt open source will probably mean that every other model that doesn't want to cheat will scrape that and accidentally cheat in the next models.

Re: Day 1 of ARC-AGI-3

#33
post #10
post #2

Note that this uses a harness so it doesn't qualify for the official ARC-AGI-3 leaderboard According to the authors the harness isn't ARC-AGI specific though https://x.com/agenticasdk/status/2037335806264971461

Doesn't the chat version of chatgpt or gemini also have interleaved tool calls, so do those also count as with harnesses?

Harness is fine. I think people here are arguing what provided here to take the test is not harness.

Re: Day 1 of ARC-AGI-3

#34
we constantly underestimate the power of inference scaffolding. I have seen it in all domains: coding, ASR, ARC-AGI benchmarks you name it. Scaffolding can do a lot! And post-training too. I am confident our currently pre-trained models can beat this benchmark over 80% with the right post-training and scaffolding. That being said I don't think ARC-AGI proves much. It is not a useful task at all in the wild. it is just a game; a strange and confusing one. For me this is just a pointless pseudo-academic exercise. Good to have, but by no means measures intelligence and even less utility of a model.

Re: Day 1 of ARC-AGI-3

#35
post #8
post #2

Note that this uses a harness so it doesn't qualify for the official ARC-AGI-3 leaderboard According to the authors the harness isn't ARC-AGI specific though https://x.com/agenticasdk/status/2037335806264971461

> this uses a harness This seems like an arbitrary restriction. Tool-use requires a harness, and their whitepaper never defines exactly what counts as valid.

It isn't arbitrary. They want measure the capability of the general LLM

Re: Day 1 of ARC-AGI-3

#36

Earlier quoted context omitted.

All traffic is monitored, all signal sources are eventually incorporated into the training set in one way or another. The person you're responding to is correct, even a single API call to any AI provider is sufficient to discount future results from the same provider.

You live in a conspiracy world. Those AI providers don't update the models that fast. You can try ask them solve ARC-AGI-3 without harness and see them struggle as yesterday yourself.

Which part is the conspiracy? Be as concrete as possible.

Re: Day 1 of ARC-AGI-3

#37
post #2

Note that this uses a harness so it doesn't qualify for the official ARC-AGI-3 leaderboard According to the authors the harness isn't ARC-AGI specific though https://x.com/agenticasdk/status/2037335806264971461

It is 100% ARC-AGI-3 specific though, just read through the prompts https://github.com/symbolica-ai/ARC-AGI-3-Agents/blob/symbol...

this is so disingenuous on symbolica's part. these insincere announcements just make it harder for genuine attempts and novel ideas

Re: Day 1 of ARC-AGI-3

#38
post #26

https://en.wikipedia.org/wiki/Goodhart's_law > Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.

[deleted]

Re: Day 1 of ARC-AGI-3

#39
The fact that this was on the set of training problems with a custom harness basically makes the headline a lie.

What if you give opus the same harness? Do people even care about meaningful comparisons any more or is it all just “numbers go up”

Re: Day 1 of ARC-AGI-3

#40

we constantly underestimate the power of inference scaffolding. I have seen it in all domains: coding, ASR, ARC-AGI benchmarks you name it. Scaffolding can do a lot! And post-training too. I am confident our currently pre-trained models can beat this benchmark over 80% with the right post-training and scaffolding. That being said I don't think ARC-AGI proves much. It is not a useful task at all in the wild. it is jus…

what exactly does scaffolding mean in this context? genuine question
Post reply on HN