Earlier quoted context omitted.
If they use Python to fill the gap, and the end user doesn’t have to know or care, is it unfair to assess this as progress and attribute the progress to the _system_? OK, the core technology that is the language model still can’t math as well as you’d hope, but how about the end result users see from the system when they interface with it? “Did you know humans are better at flying today than they were a thousand year…
You are correct, the frameworks around it have improved. In that regard, my assessment is unfair: I only judge the underlying technology and what is sold by the sota providers, with the premise of what it's like when you start fresh. You can achieve a lot by coding around the issues, but that's kinda against the point of 'AI', is it?
ARC-AGI Leaderboard
91–100 of 156 posts
Re: ARC-AGI Leaderboard
#92Why Anthropic models are always leapfrogging these benchmarks, but in real life work I do feel like after 3 weeks I am back to Claude Opus 4.5? (regardless of the model I use, Fable was exception for 1 day when it was released)
It's called frog boiling. We get used to the new level of intelligence so fast, any deviation feels like going back to the stone age. If you don't believe me, create something complex with Opus 5 and then with Opus 4.5, and notice the difference.
Re: ARC-AGI Leaderboard
#93Re: ARC-AGI Leaderboard
#94Appears to be benchmaxxing https://x.com/quietnning/status/2080786711861407883
I guess there is no way this can happen without benchmark being part of the training data??
Re: ARC-AGI Leaderboard
#95Why Anthropic models are always leapfrogging these benchmarks, but in real life work I do feel like after 3 weeks I am back to Claude Opus 4.5? (regardless of the model I use, Fable was exception for 1 day when it was released)
5 seems incredibly smart to me in my conversations today about some pretty niche ideas in.longitudinal modeling. It.felt.like a big step up.from 4.8, to me
Re: ARC-AGI Leaderboard
#96test Codex, not Sol. test Claude code, not Opus
Re: ARC-AGI Leaderboard
#97Re: ARC-AGI Leaderboard
#98Re: ARC-AGI Leaderboard
#99Appears to be benchmaxxing https://x.com/quietnning/status/2080786711861407883