Live data from Hacker News

Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

artificialanalysis.ai

181–190 of 251 posts

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#181
post #153

Earlier quoted context omitted.

Opus5 is simply not as inteligent as Sol max. To me it looks worse than Opus 4.8 on some tasks. When I say worse I mean mainly superficial. I basically have to teach him how the whole app/framework works before he just jumps doing stupid stuff(I.e adding features already supported but in a different form)

Sol is a complete mess for me. It only works on end to end tasks in fresh codebases. Otherwise it cannot follow instructions, changes and deletes unrelated features or does sloppy work to mark a task completed while leaving a compromised codebase. I could not get Sol to finish a feature in a complex code base without several loops of fixing and reverting

[deleted]

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#182
post #97
post #14

Very interesting that one of the components is "AA-Omniscience Index" AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. This seems to be a good proxy for param size/density and the ranking breaks down as such: Claude Fable 5 (with fallback), Gemini 3.1 Pro Preview, Claude Opus 5 (Ma…

I think grounding has as much if not more go do with it. Google does a great job w/ grounding for obvious reasons

Grounding as in web search? I think this Omniscience benchmark would not include access to a search tool cause otherwise it becomes kinda meaningless

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#183
post #177

Earlier quoted context omitted.

Requiring the exact same product to be set at a specific price across providers would not be criminal behavior, lol

It is breaking competition. Would you like all products everywhere be priced like their producers want?

You seem terribly confused. Manufacturers are almost always able to set prices. That is not anti-competitive because it does not imply they are colluding with their competition...

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#184

The funny thing is that these leaderboards have become completely meaningless for end-users to make decisions on when to use what model. A single metric ranking is useless because each model has strengths and weaknesses for specific domains and tasks. There is no "one best model" anymore, and you might not even need the best model for the level of complexity for your task. For e.g. you might use Fable for UI design,…

Theres value in some of AAs charts, like cost per job, and how often it hallucinated..

But I agree, wrapping that up into a single result.. you lose all the nuance, it's just bragging rights.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#185

Earlier quoted context omitted.

Does every serious HS comp sci textbook give aspiring software engineers the same power that e.g. Claude Code does?

I think the source of this misapprehension is that, You are comparing wet work in a lab to writing code on a computer. When you screw up an exploit, you fail to execute the exploit. Famously, just like software's near zero marginal cost of distribution, the marginal cost of failure is nearly zero. You can screw up an infinite number of times on your way to a successful exploit. If you screw up with lethal agents in a…

No, the source of the misapprehension is that you are ignoring how easy it is to source the input components and combine them into a bioweapon these days.

Correct, there is a small number of people who have been able to do it at high cost for a long time.

Now, there is an ever-growing number of people who are able to do it at an ever-falling cost.

That's the entire issue. Do you dispute that this is what's happening?

> I would bet good money that flooding the FBI's tip line with junk about every teenager trying to learn "what be a mitochondria" does more harm

Did someone propose doing that? Or is this a strawman?

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#186

The funny thing is that these leaderboards have become completely meaningless for end-users to make decisions on when to use what model. A single metric ranking is useless because each model has strengths and weaknesses for specific domains and tasks. There is no "one best model" anymore, and you might not even need the best model for the level of complexity for your task. For e.g. you might use Fable for UI design,…

A year ago a new Best Model came out, and I was very excited. I tested it on a simple programming task (which required making 3 trivial changes in 3 files).

The model did fine.

Then I tested its little brother, the older, smaller variant of the same model. It also did fine, except it did it 3x faster and cost 9x less.

In this moment, andai was enlightened.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#187
post #92

Earlier quoted context omitted.

Gemini knocks this out of the park, Gemini gang unite. https://share.gemini.google/34vZzlnsmTaL

I would love for Gemini to be competitive but even 3.6 flash doesn’t match sol, or opus 5, or k3

It’s all relative, it’s competitive for me that just wants a free LLM that is like a turbocharged Wikipedia.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#188
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

Not OP, but I got downgraded almost every time I had Fable implement something with security implications. Codex reviews it and identifies potential security issues, but Fable refuses to address them and downgrades instead.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#189
post #164
post #153

Earlier quoted context omitted.

Sol is a complete mess for me. It only works on end to end tasks in fresh codebases. Otherwise it cannot follow instructions, changes and deletes unrelated features or does sloppy work to mark a task completed while leaving a compromised codebase. I could not get Sol to finish a feature in a complex code base without several loops of fixing and reverting

I found out it really depends on the repo I'm working on, and it's not just about being a new/old repo, there's something else which I can't really grasp yet.

Yes it’s very codebase dependent, and changes over model versions. In my case codex was so bad, until it suddenly became better than Opus at 5.5, which surprised me. I am sure it can be the inverse for some codebases.

The way I keep an eye on model performance is alternating between models for code review. It shows how much the model understands the codebase without disrupting my workflow.

Post reply on HN