Live data from Hacker News

ARC-AGI Leaderboard

arcprize.org

41–50 of 156 posts

Re: ARC-AGI Leaderboard

#42

It's actually crazy to see the difference between opus 5 and the next best model on ARC AGI 3 when you actually look at the ARC AGI problems

Why? 30% is passing the first two problems only, which are really very simple.

Re: ARC-AGI Leaderboard

#43

Earlier quoted context omitted.

For frontier models, not local. https://schema-harness.github.io/

Yes, saw that. They haven't yet released any code. Until they do, treat it with a huuuge grain of salt. In fact treat any 99% result in ML with a huge grain of salt.

If you stop and think about the problem it really is quite simple. Just need to build a graph of the game state and then run A* to get to the end.

Re: ARC-AGI Leaderboard

#44
post #40

> Only systems which required less than $10,000 to run are shown. (Notes[1]) Am I lost or are their many models on this ranking (Opus 5 included) that clear this?

Many models are much cheaper through their subscriptions' included usage. That could be what's happening here.

Claude gives you something like $5000 of tokens on a $200 plan.

Re: ARC-AGI Leaderboard

#45
post #34

Earlier quoted context omitted.

I've worked with these systems for four years now and they have not meaningfully improved in that time frame. We still have: - statistical correlation between two things will always cause one thing to lead to the other, no matter how much you prompt it to not have that connection (to be expected with a stochastic system) - Math completely fails in longer contexts - "thinking" token generation being on the correct tra…

> - Math completely fails in longer contexts Not sure what longer contexts we're talking about but didn't we have an old math problem optimized, which even the LLM itself was surprised about, just a week ago? Something which wasn't possible 6 months ago.

I mean calculations, not mathematical proofs

Re: ARC-AGI Leaderboard

#46
post #34

Why Anthropic models are always leapfrogging these benchmarks, but in real life work I do feel like after 3 weeks I am back to Claude Opus 4.5? (regardless of the model I use, Fable was exception for 1 day when it was released)

I've worked with these systems for four years now and they have not meaningfully improved in that time frame. We still have: - statistical correlation between two things will always cause one thing to lead to the other, no matter how much you prompt it to not have that connection (to be expected with a stochastic system) - Math completely fails in longer contexts - "thinking" token generation being on the correct tra…

Like how toddlers’ skills don’t meaningfully improve on infants’, because either could wake up in a wet bed.

Re: ARC-AGI Leaderboard

#47
post #39
post #34

Earlier quoted context omitted.

I've worked with these systems for four years now and they have not meaningfully improved in that time frame. We still have: - statistical correlation between two things will always cause one thing to lead to the other, no matter how much you prompt it to not have that connection (to be expected with a stochastic system) - Math completely fails in longer contexts - "thinking" token generation being on the correct tra…

> I've worked with these systems for four years now and they have not meaningfully improved in that time frame. Not meaningfully improved?! Four years ago was gpt *3.5*! ChatGPT hadn’t been released!

Yes! Impressive, isn't it? I see how it has improved for some minor points, that the big models can cover more finetuning ground, but my big gripes are still the same - you could do the same back then with multiple models and more targeted finetuning.

Re: ARC-AGI Leaderboard

#48
post #46
post #34

Earlier quoted context omitted.

I've worked with these systems for four years now and they have not meaningfully improved in that time frame. We still have: - statistical correlation between two things will always cause one thing to lead to the other, no matter how much you prompt it to not have that connection (to be expected with a stochastic system) - Math completely fails in longer contexts - "thinking" token generation being on the correct tra…

Like how toddlers’ skills don’t meaningfully improve on infants’, because either could wake up in a wet bed.

Let us be more clear: there is no structural jump, no architectural overcoming of the original fault.

(Edit: and on a similar point, structural properties such as having static ntetworks, as opposed to continuously learning and improving architectures (such as us), will reveal that there is still road ahead.)

Re: ARC-AGI Leaderboard

#49
post #47
post #39

Earlier quoted context omitted.

> I've worked with these systems for four years now and they have not meaningfully improved in that time frame. Not meaningfully improved?! Four years ago was gpt *3.5*! ChatGPT hadn’t been released!

Yes! Impressive, isn't it? I see how it has improved for some minor points, that the big models can cover more finetuning ground, but my big gripes are still the same - you could do the same back then with multiple models and more targeted finetuning.

> you could do the same back then with multiple models and more targeted finetuning

Are you one of those anonymous billionaires as if you did this a few years ago, you would've been famous and rich.

Re: ARC-AGI Leaderboard

#50
post #32
post #21

Earlier quoted context omitted.

How do they handle these assurances? Personally I have zero trust in the AI companies not trying to use this data to get ahead in the game, and short of sharing the weights and harness so that the benchmarkers can run the models themselves, I don't see a satisfactory solution with this mindset.

OpenAI's Zero Data Retention claim held up in court. They were unable to produce prompts and outputs because they were never retained. I believe that is only available through Enterprise API for both Anthropic and OpenAI.

Is there a distinction we can independently assess between never retained and deleted or hidden?

Asking especially given CEO’s track record https://news.ycombinator.com/item?id=47659135

Post reply on HN