Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
211–220 of 251 posts
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#212Earlier quoted context omitted.
This sort of rhetoric appears for every benchmark, and somehow it always rises to the top. A few days ago there was a Geekbench 7 submission on here ( https://news.ycombinator.com/item?id=49025812 ), and again the top comment was someone dismissing it, using the classic "but I want a benchmark specifically for exactly the thing I do" perfect-is-the-enemy-of-good nonsense. I, one of those end users, absolutely use the…
Firstly, I don't have many issues with benchmarks per se. But I do have issues with leaderboards. And the AA index is touted by lots of people to argue that X model is better than Y, which I find inaccurate. > I mean, it's telling that your gamut of examples are three models that are within spitting distance of each other on the broadest benchmarks. This is kind of my point. The benchmarks say they are splitting dist…
It's a shit point, then. And absolutely no one said they were "equivalent", and again you're doing the rhetorical "it isn't perfect and absolutely comprehensive for every possible scenario, therefore it is "completely meaningless". Again, you chose that absurd terminology, rather than for instance "doesn't tell the whole story".
Again, you chose three models for your example at the very tops of the leaderboards. The SOTA models. Which kind of means that the leaderboards actually mean an incredible amount, no?
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#213Earlier quoted context omitted.
A year ago a new Best Model came out, and I was very excited. I tested it on a simple programming task (which required making 3 trivial changes in 3 files). The model did fine. Then I tested its little brother, the older, smaller variant of the same model. It also did fine, except it did it 3x faster and cost 9x less. In this moment, andai was enlightened.
is there a benchmark that uses prices or speed as one of the axis, in addition to accuracy? Best could mean different things to different people.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#214The funny thing is that these leaderboards have become completely meaningless for end-users to make decisions on when to use what model. A single metric ranking is useless because each model has strengths and weaknesses for specific domains and tasks. There is no "one best model" anymore, and you might not even need the best model for the level of complexity for your task. For e.g. you might use Fable for UI design,…
A year ago a new Best Model came out, and I was very excited. I tested it on a simple programming task (which required making 3 trivial changes in 3 files). The model did fine. Then I tested its little brother, the older, smaller variant of the same model. It also did fine, except it did it 3x faster and cost 9x less. In this moment, andai was enlightened.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#215Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#216So if you're in Google leadership, you sleep in the office, right? Not merely because you have a ton of work but also because you're deeply ashamed to be seen in public.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#217Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#218So if you're in Google leadership, you sleep in the office, right? Not merely because you have a ton of work but also because you're deeply ashamed to be seen in public.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#219Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#220The funny thing is that these leaderboards have become completely meaningless for end-users to make decisions on when to use what model. A single metric ranking is useless because each model has strengths and weaknesses for specific domains and tasks. There is no "one best model" anymore, and you might not even need the best model for the level of complexity for your task. For e.g. you might use Fable for UI design,…
There's some kernel of truth to what you are saying, but hyperboles like this just aren't accurate. All statistics lie but its better than being blind... What your post really says is that benchmarks only show an average over multiple tasks. Yes, obviously, the point of a statistic is to summarize.