Appears to be benchmaxxing https://x.com/quietnning/status/2080786711861407883
"Opus 5 states the hidden rules before its first action, then plays a byte-identical optimal solution in 5/5 seeds at temperature 1.0. Zero exploration." I guess there is no way this can happen without benchmark being part of the training data??
ARC-AGI Leaderboard
101–110 of 156 posts
Re: ARC-AGI Leaderboard
#102Appears to be benchmaxxing https://x.com/quietnning/status/2080786711861407883
"Opus 5 states the hidden rules before its first action, then plays a byte-identical optimal solution in 5/5 seeds at temperature 1.0. Zero exploration." I guess there is no way this can happen without benchmark being part of the training data??
Re: ARC-AGI Leaderboard
#103Appears to be benchmaxxing https://x.com/quietnning/status/2080786711861407883
Re: ARC-AGI Leaderboard
#104It's actually crazy to see the difference between opus 5 and the next best model on ARC AGI 3 when you actually look at the ARC AGI problems
Why? 30% is passing the first two problems only, which are really very simple.
Re: ARC-AGI Leaderboard
#105this is not a good measure of current model capability. we need to test agents in harnesses, not models with a single prompt test Codex, not Sol. test Claude code, not Opus
Re: ARC-AGI Leaderboard
#106Appears to be benchmaxxing https://x.com/quietnning/status/2080786711861407883
I'm shocked, astonished even, that enterprises on which trillions of dollars are being poured would consider cheating on marketing benchmarks.
It's more likely that the training data was contaminated with the benchmark data.
Re: ARC-AGI Leaderboard
#107but I kinda wish I could select level... I accidentally pressed redirect button and when I came back I was once again shown level 1, all progress lost :(
Re: ARC-AGI Leaderboard
#108Earlier quoted context omitted.
I'm shocked, astonished even, that enterprises on which trillions of dollars are being poured would consider cheating on marketing benchmarks.
Meh, I doubt it was intentional. Deliberate benchmaxxing is incredibly damaging to credibility once it's discovered (see what happened to Meta with LlaMa 4). It's more likely that the training data was contaminated with the benchmark data.
They maybe have not intentionally benchmaxxed, but they certainly know that's what happened .
Re: ARC-AGI Leaderboard
#109Why Anthropic models are always leapfrogging these benchmarks, but in real life work I do feel like after 3 weeks I am back to Claude Opus 4.5? (regardless of the model I use, Fable was exception for 1 day when it was released)
For the first few weeks, they would actually provide those 660 minutes to establish trust in their cards. But after a while, they would quietly start reducing the number of minutes on subsequent top-ups—say, from 660 minutes down to only 300. They wouldn’t do this for every card, so it was difficult to prove. Instead, they relied on averages across their customer base to make the economics work.
Lately, I’ve found myself wondering whether something similar may be happening with frontier AI models. Companies launch with an exceptionally strong model and generous compute limits to build adoption. Once the model is established as a market leader, the incentives change, and users may start perceiving the service as becoming more constrained or less capable over time.
I don’t have evidence that this is what’s happening with Anthropic—or with any other AI company. It’s simply a pattern that the current situation reminds me of.
Re: ARC-AGI Leaderboard
#110Earlier quoted context omitted.
"Opus 5 states the hidden rules before its first action, then plays a byte-identical optimal solution in 5/5 seeds at temperature 1.0. Zero exploration." I guess there is no way this can happen without benchmark being part of the training data??
What seems to implied is that some of his hold out testing suite includes simple/common tests that are out in the wild, and for those opus went straight to a memorised solution .