Live data from Hacker News

Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

artificialanalysis.ai

211–220 of 251 posts

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#211
New respect for Meta Muse Spark. It seems to sit at a lot of sweet spots in the leader board. It’s not the best at anything in particular, but it balances cost and performance quite well. I’m also curious where Poolside Laguna S would sit; it’s not included. I’m personally very interested in cost effective models that still perform well.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#212

Earlier quoted context omitted.

This sort of rhetoric appears for every benchmark, and somehow it always rises to the top. A few days ago there was a Geekbench 7 submission on here ( https://news.ycombinator.com/item?id=49025812 ), and again the top comment was someone dismissing it, using the classic "but I want a benchmark specifically for exactly the thing I do" perfect-is-the-enemy-of-good nonsense. I, one of those end users, absolutely use the…

Firstly, I don't have many issues with benchmarks per se. But I do have issues with leaderboards. And the AA index is touted by lots of people to argue that X model is better than Y, which I find inaccurate. > I mean, it's telling that your gamut of examples are three models that are within spitting distance of each other on the broadest benchmarks. This is kind of my point. The benchmarks say they are splitting dist…

>This is kind of my point.

It's a shit point, then. And absolutely no one said they were "equivalent", and again you're doing the rhetorical "it isn't perfect and absolutely comprehensive for every possible scenario, therefore it is "completely meaningless". Again, you chose that absurd terminology, rather than for instance "doesn't tell the whole story".

Again, you chose three models for your example at the very tops of the leaderboards. The SOTA models. Which kind of means that the leaderboards actually mean an incredible amount, no?

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#213
post #186

Earlier quoted context omitted.

A year ago a new Best Model came out, and I was very excited. I tested it on a simple programming task (which required making 3 trivial changes in 3 files). The model did fine. Then I tested its little brother, the older, smaller variant of the same model. It also did fine, except it did it 3x faster and cost 9x less. In this moment, andai was enlightened.

is there a benchmark that uses prices or speed as one of the axis, in addition to accuracy? Best could mean different things to different people.

Yes same website https://artificialanalysis.ai/models#intelligence-comparison... but they don't have graphs for the individual benchmarks sadly.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#214
post #186

The funny thing is that these leaderboards have become completely meaningless for end-users to make decisions on when to use what model. A single metric ranking is useless because each model has strengths and weaknesses for specific domains and tasks. There is no "one best model" anymore, and you might not even need the best model for the level of complexity for your task. For e.g. you might use Fable for UI design,…

A year ago a new Best Model came out, and I was very excited. I tested it on a simple programming task (which required making 3 trivial changes in 3 files). The model did fine. Then I tested its little brother, the older, smaller variant of the same model. It also did fine, except it did it 3x faster and cost 9x less. In this moment, andai was enlightened.

Sure, that's true until you hit a difficult problem where the smaller models thrash endlessly whereas its big sibling can solve it with one prompt.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#216
post #190

So if you're in Google leadership, you sleep in the office, right? Not merely because you have a ton of work but also because you're deeply ashamed to be seen in public.

If you're Demis, at least, you sleep fine because you were personally an early investor in Anthropic.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#218
post #190

So if you're in Google leadership, you sleep in the office, right? Not merely because you have a ton of work but also because you're deeply ashamed to be seen in public.

Is Google trying to compete with OpenAI and Anthropic re: maximally intelligent models? Google seems to be the only one of the three that doesn't pray and self flagellate at the altar of AGI.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#220

The funny thing is that these leaderboards have become completely meaningless for end-users to make decisions on when to use what model. A single metric ranking is useless because each model has strengths and weaknesses for specific domains and tasks. There is no "one best model" anymore, and you might not even need the best model for the level of complexity for your task. For e.g. you might use Fable for UI design,…

"completely meaningless" "the only purpose"

There's some kernel of truth to what you are saying, but hyperboles like this just aren't accurate. All statistics lie but its better than being blind... What your post really says is that benchmarks only show an average over multiple tasks. Yes, obviously, the point of a statistic is to summarize.

Post reply on HN