Live data from Hacker News

Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

artificialanalysis.ai

201–210 of 251 posts

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#201

Earlier quoted context omitted.

Opus5 is simply not as inteligent as Sol max. To me it looks worse than Opus 4.8 on some tasks. When I say worse I mean mainly superficial. I basically have to teach him how the whole app/framework works before he just jumps doing stupid stuff(I.e adding features already supported but in a different form)

Opus 5 hasn't been available for that long - long enough for benchmarks, but not really use and develop a subjective view on

I have spent about 6 hours with it now and it is absurd to say it worse than 4.8. It is wonderful.

To me, there is this strange critique that seems to always happen now with a new model. Like shitting on the model for entertainment purposes seems more interesting to many than actually using the model.

It reminds me of looking up a new music album on youtube that has very few views with a review above it by The Needle Drop shitting on the album will have a few hundred thousands views.

I would rather just listen to the album and judge for myself.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#202

The funny thing is that these leaderboards have become completely meaningless for end-users to make decisions on when to use what model. A single metric ranking is useless because each model has strengths and weaknesses for specific domains and tasks. There is no "one best model" anymore, and you might not even need the best model for the level of complexity for your task. For e.g. you might use Fable for UI design,…

If you click on the link you will see that it's not "one single metric" there is literally all the metrics so you can make an informed decision

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#203
post #186

The funny thing is that these leaderboards have become completely meaningless for end-users to make decisions on when to use what model. A single metric ranking is useless because each model has strengths and weaknesses for specific domains and tasks. There is no "one best model" anymore, and you might not even need the best model for the level of complexity for your task. For e.g. you might use Fable for UI design,…

A year ago a new Best Model came out, and I was very excited. I tested it on a simple programming task (which required making 3 trivial changes in 3 files). The model did fine. Then I tested its little brother, the older, smaller variant of the same model. It also did fine, except it did it 3x faster and cost 9x less. In this moment, andai was enlightened.

is there a benchmark that uses prices or speed as one of the axis, in addition to accuracy?

Best could mean different things to different people.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#204

Earlier quoted context omitted.

I think the source of this misapprehension is that, You are comparing wet work in a lab to writing code on a computer. When you screw up an exploit, you fail to execute the exploit. Famously, just like software's near zero marginal cost of distribution, the marginal cost of failure is nearly zero. You can screw up an infinite number of times on your way to a successful exploit. If you screw up with lethal agents in a…

No, the source of the misapprehension is that you are ignoring how easy it is to source the input components and combine them into a bioweapon these days. Correct, there is a small number of people who have been able to do it at high cost for a long time. Now, there is an ever-growing number of people who are able to do it at an ever-falling cost. That's the entire issue. Do you dispute that this is what's happening?…

> No, the source of the misapprehension is that you are ignoring how easy it is to source the input components and combine them into a bioweapon these days.

Citation needed. Where are these home biolabs? Why haven't any leaked yet like home meth labs?

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#205
post #190

So if you're in Google leadership, you sleep in the office, right? Not merely because you have a ton of work but also because you're deeply ashamed to be seen in public.

They are tied for first using the Google-proof question and answer benchmark: https://artificialanalysis.ai/evaluations/gpqa-diamond

Maybe that's their only goal?

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#206
post #190

So if you're in Google leadership, you sleep in the office, right? Not merely because you have a ton of work but also because you're deeply ashamed to be seen in public.

Gemini models are at or near the top in several categories, though, so I'm not sure the takeaway is that they're shamefully far behind.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#207

Why do I find it dummer/even more superficial than opus 4.8? It just continued a session and I had to stop it because it become obviously “lost”

Massive performance degradation is expected when continuing a session with a different model

Is it? Is there a paper that covers this?

I would've thought it should be mostly seamless, since it's being fed the entire conversation all along anyway.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#208

The funny thing is that these leaderboards have become completely meaningless for end-users to make decisions on when to use what model. A single metric ranking is useless because each model has strengths and weaknesses for specific domains and tasks. There is no "one best model" anymore, and you might not even need the best model for the level of complexity for your task. For e.g. you might use Fable for UI design,…

This sort of rhetoric appears for every benchmark, and somehow it always rises to the top. A few days ago there was a Geekbench 7 submission on here (https://news.ycombinator.com/item?id=49025812), and again the top comment was someone dismissing it, using the classic "but I want a benchmark specifically for exactly the thing I do" perfect-is-the-enemy-of-good nonsense.

I, one of those end users, absolutely use these benchmarks as heavy input considerations. Indeed, the vast majority of people do. "Completely meaningless" is just nonsense, of course, and while it doesn't perfectly map to every use, there is a pretty good correlation with suitability for specific tasks.

I mean, it's telling that your gamut of examples are three models that are within spitting distance of each other on the broadest benchmarks.

Not to mention that the linked page includes a pretty broad list of specialization benchmarks.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#209

The funny thing is that these leaderboards have become completely meaningless for end-users to make decisions on when to use what model. A single metric ranking is useless because each model has strengths and weaknesses for specific domains and tasks. There is no "one best model" anymore, and you might not even need the best model for the level of complexity for your task. For e.g. you might use Fable for UI design,…

This sort of rhetoric appears for every benchmark, and somehow it always rises to the top. A few days ago there was a Geekbench 7 submission on here ( https://news.ycombinator.com/item?id=49025812 ), and again the top comment was someone dismissing it, using the classic "but I want a benchmark specifically for exactly the thing I do" perfect-is-the-enemy-of-good nonsense. I, one of those end users, absolutely use the…

Firstly, I don't have many issues with benchmarks per se. But I do have issues with leaderboards. And the AA index is touted by lots of people to argue that X model is better than Y, which I find inaccurate.

> I mean, it's telling that your gamut of examples are three models that are within spitting distance of each other on the broadest benchmarks.

This is kind of my point. The benchmarks say they are splitting distance, but they actually vary wildly in performance for specific tasks, so they are in fact not equivalent.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#210

Earlier quoted context omitted.

No, the source of the misapprehension is that you are ignoring how easy it is to source the input components and combine them into a bioweapon these days. Correct, there is a small number of people who have been able to do it at high cost for a long time. Now, there is an ever-growing number of people who are able to do it at an ever-falling cost. That's the entire issue. Do you dispute that this is what's happening?…

> No, the source of the misapprehension is that you are ignoring how easy it is to source the input components and combine them into a bioweapon these days. Citation needed. Where are these home biolabs? Why haven't any leaked yet like home meth labs?

Who said anything about a home biolab? Are you thinking a possible solution is just to block the terrorists, irresponsible corporations, or evil governments from LLMs? Obviously not.

In any case, the "home biolab" required to do this stuff gets smaller and more accessible every day. Biochemistry, like virtually every other complex procedural field, has become heavily outsourced. You can literally order genetic fragments even of known pathogens on the Internet, shipped to your door. There should obviously be much more aggressive restrictions on manufacturing known pathogen fragments, but 1) every money-hungry lab would need to volunteer to participate, and 2) it's totally unclear how they'd detect novel pathogen fragments that unsafe AI would be happy to help predict a couple thousand of.

Today, composing that into a working virus might require an undergrad biochem education, a hundred grand, and a bunch of patience (and risk), but that describes millions of people. As GP pointed out, it wasn't too long ago the group of people with this capability was fewer than a dozen individuals on the planet. And as is obvious, we are trending in one direction. We are not trending the other direction.

Post reply on HN