Live data from Hacker News

ARC-AGI Leaderboard

arcprize.org

91–100 of 156 posts

Re: ARC-AGI Leaderboard

#91
post #73
post #51

Earlier quoted context omitted.

If they use Python to fill the gap, and the end user doesn’t have to know or care, is it unfair to assess this as progress and attribute the progress to the _system_? OK, the core technology that is the language model still can’t math as well as you’d hope, but how about the end result users see from the system when they interface with it? “Did you know humans are better at flying today than they were a thousand year…

You are correct, the frameworks around it have improved. In that regard, my assessment is unfair: I only judge the underlying technology and what is sold by the sota providers, with the premise of what it's like when you start fresh. You can achieve a lot by coding around the issues, but that's kinda against the point of 'AI', is it?

No, its capabilities with a harness are what we are interested in. Your assessment is only relevant to benchmarking, not practical value.

Re: ARC-AGI Leaderboard

#92
post #28

Why Anthropic models are always leapfrogging these benchmarks, but in real life work I do feel like after 3 weeks I am back to Claude Opus 4.5? (regardless of the model I use, Fable was exception for 1 day when it was released)

It's called frog boiling. We get used to the new level of intelligence so fast, any deviation feels like going back to the stone age. If you don't believe me, create something complex with Opus 5 and then with Opus 4.5, and notice the difference.

The actual term for this is hedonic adaptation.

Re: ARC-AGI Leaderboard

#94

Appears to be benchmaxxing https://x.com/quietnning/status/2080786711861407883

"Opus 5 states the hidden rules before its first action, then plays a byte-identical optimal solution in 5/5 seeds at temperature 1.0. Zero exploration."

I guess there is no way this can happen without benchmark being part of the training data??

Re: ARC-AGI Leaderboard

#95

Why Anthropic models are always leapfrogging these benchmarks, but in real life work I do feel like after 3 weeks I am back to Claude Opus 4.5? (regardless of the model I use, Fable was exception for 1 day when it was released)

5 seems incredibly smart to me in my conversations today about some pretty niche ideas in.longitudinal modeling. It.felt.like a big step up.from 4.8, to me

5 felt both smarter than me and dumber in some ways - it gets stuck to its original ideas. I had never seen a model harder to talk into changing its initial opinions. it continuously hedges.

Re: ARC-AGI Leaderboard

#99

Appears to be benchmaxxing https://x.com/quietnning/status/2080786711861407883

I was just thinking they need to mark each model per benchmark as "model released before the benchmark was released" and "model released after the benchmark was released".

Re: ARC-AGI Leaderboard

#100

I have a suspicion that they are just trained on puzzles by now

It's like we've come full circle: First people practiced L33t3cod3 problems for interviews Then people built AIs to build software And now the AIs are studying L33t3cod3 problems

Why the 3 rather than e?
Post reply on HN