Live data from Hacker News

ARC-AGI Leaderboard

arcprize.org

101–110 of 156 posts

Re: ARC-AGI Leaderboard

#101

Appears to be benchmaxxing https://x.com/quietnning/status/2080786711861407883

"Opus 5 states the hidden rules before its first action, then plays a byte-identical optimal solution in 5/5 seeds at temperature 1.0. Zero exploration." I guess there is no way this can happen without benchmark being part of the training data??

What seems to implied is that some of his hold out testing suite includes simple/common tests that are out in the wild, and for those opus went straight to a memorised solution .

Re: ARC-AGI Leaderboard

#102

Appears to be benchmaxxing https://x.com/quietnning/status/2080786711861407883

"Opus 5 states the hidden rules before its first action, then plays a byte-identical optimal solution in 5/5 seeds at temperature 1.0. Zero exploration." I guess there is no way this can happen without benchmark being part of the training data??

It may read information about the benchmark, such as on blog post, or Twitter feeds (example OP) without active cheat

Re: ARC-AGI Leaderboard

#104
post #42

It's actually crazy to see the difference between opus 5 and the next best model on ARC AGI 3 when you actually look at the ARC AGI problems

Why? 30% is passing the first two problems only, which are really very simple.

Huh. How do things end up with scores like 30.2% (and results between 0% and 1%) if it's that low resolution?

Re: ARC-AGI Leaderboard

#106

Appears to be benchmaxxing https://x.com/quietnning/status/2080786711861407883

I'm shocked, astonished even, that enterprises on which trillions of dollars are being poured would consider cheating on marketing benchmarks.

Meh, I doubt it was intentional. Deliberate benchmaxxing is incredibly damaging to credibility once it's discovered (see what happened to Meta with LlaMa 4).

It's more likely that the training data was contaminated with the benchmark data.

Re: ARC-AGI Leaderboard

#107
games are great (as for a human)

but I kinda wish I could select level... I accidentally pressed redirect button and when I came back I was once again shown level 1, all progress lost :(

Re: ARC-AGI Leaderboard

#108

Earlier quoted context omitted.

I'm shocked, astonished even, that enterprises on which trillions of dollars are being poured would consider cheating on marketing benchmarks.

Meh, I doubt it was intentional. Deliberate benchmaxxing is incredibly damaging to credibility once it's discovered (see what happened to Meta with LlaMa 4). It's more likely that the training data was contaminated with the benchmark data.

You really think they saw the jump in arc-agi-3 (which they reported in their official card), and didn't even bother to check?

They maybe have not intentionally benchmaxxed, but they certainly know that's what happened .

Re: ARC-AGI Leaderboard

#109

Why Anthropic models are always leapfrogging these benchmarks, but in real life work I do feel like after 3 weeks I am back to Claude Opus 4.5? (regardless of the model I use, Fable was exception for 1 day when it was released)

Some 20 years ago, the telecommunications sector in Germany was liberalized. Many telephone card providers entered what had previously been a barely competitive market. They advertised their products with aggressive claims like: “Buy our €10 top-up card and get 660 minutes to destination X.”

For the first few weeks, they would actually provide those 660 minutes to establish trust in their cards. But after a while, they would quietly start reducing the number of minutes on subsequent top-ups—say, from 660 minutes down to only 300. They wouldn’t do this for every card, so it was difficult to prove. Instead, they relied on averages across their customer base to make the economics work.

Lately, I’ve found myself wondering whether something similar may be happening with frontier AI models. Companies launch with an exceptionally strong model and generous compute limits to build adoption. Once the model is established as a market leader, the incentives change, and users may start perceiving the service as becoming more constrained or less capable over time.

I don’t have evidence that this is what’s happening with Anthropic—or with any other AI company. It’s simply a pattern that the current situation reminds me of.

Re: ARC-AGI Leaderboard

#110

Earlier quoted context omitted.

"Opus 5 states the hidden rules before its first action, then plays a byte-identical optimal solution in 5/5 seeds at temperature 1.0. Zero exploration." I guess there is no way this can happen without benchmark being part of the training data??

What seems to implied is that some of his hold out testing suite includes simple/common tests that are out in the wild, and for those opus went straight to a memorised solution .

Simple/common tests is not an explanation for why only now Opus 5 is the only model encoding the answers like this. Something like the holdout test suite being leaked or Anthropic cheating (e.g. 'accidentally' including previous hold out run data in Opus 5 training) makes a much stronger fit.
Post reply on HN