Live data from Hacker News

Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

artificialanalysis.ai

161–170 of 251 posts

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#161

Earlier quoted context omitted.

Opus5 is simply not as inteligent as Sol max. To me it looks worse than Opus 4.8 on some tasks. When I say worse I mean mainly superficial. I basically have to teach him how the whole app/framework works before he just jumps doing stupid stuff(I.e adding features already supported but in a different form)

Opus 5 hasn't been available for that long - long enough for benchmarks, but not really use and develop a subjective view on

I don’t know that individuals can really be expected to use a new model enough to develop a proper opinion though. If I try a new model, and it doesn’t seem as good as the one I’m using, I’m just going to stop using it. That’s not enough data to give anyone else a useful view on it, but it’s enough for me to make my mind up.

Especially because I’m probably trying it at work, and I can’t really justify using the company’s enterprise plan to develop my understanding of a model that I don’t think is going to be the one I use for my work.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#162
post #153

Earlier quoted context omitted.

Opus5 is simply not as inteligent as Sol max. To me it looks worse than Opus 4.8 on some tasks. When I say worse I mean mainly superficial. I basically have to teach him how the whole app/framework works before he just jumps doing stupid stuff(I.e adding features already supported but in a different form)

Sol is a complete mess for me. It only works on end to end tasks in fresh codebases. Otherwise it cannot follow instructions, changes and deletes unrelated features or does sloppy work to mark a task completed while leaving a compromised codebase. I could not get Sol to finish a feature in a complex code base without several loops of fixing and reverting

It happened to me as well but in a different direction: i.e adds non library code in a shared library. Another issue I with GPT is that it is chasing too much edge cases/security issues(I.e chasing ghosts).

However this makes it also a strong model because it fixes/solves problems that both Opus and Fable are incapable. In reviews it catches bugs that both Fable and Opus are missing to spot.

To me the “best of both worlds” is to research the problem with GPT sol, create a plan with Fable and dual review it with both Fable and GPT-SOL and implement it with GPT-SOL. You can see in the code reviews how many times both Fable and opus are sloppy and superficial while GPT-SOL just does its due diligence …I had several problems that Fable just gave up and it was GPT that helped it sort it out.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#163

Earlier quoted context omitted.

Even worse, it's offensive and kids could see it!!

If it goes on like this, American AI will eventually be able to solve the Riemann hypothesis but deny the existence of nipples.

From the late Perscheid:

https://martin-perscheid.de/image/cartoon/3212.gif

("How do you reliably put an American out-of-combat.")

But the same vibe is felt outside the USA anyway (the uk came close to that attitude with declarations from starmer earlier this year).

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#164
post #153

Earlier quoted context omitted.

Opus5 is simply not as inteligent as Sol max. To me it looks worse than Opus 4.8 on some tasks. When I say worse I mean mainly superficial. I basically have to teach him how the whole app/framework works before he just jumps doing stupid stuff(I.e adding features already supported but in a different form)

Sol is a complete mess for me. It only works on end to end tasks in fresh codebases. Otherwise it cannot follow instructions, changes and deletes unrelated features or does sloppy work to mark a task completed while leaving a compromised codebase. I could not get Sol to finish a feature in a complex code base without several loops of fixing and reverting

I found out it really depends on the repo I'm working on, and it's not just about being a new/old repo, there's something else which I can't really grasp yet.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#165
post #56

What's interesting is this: The top AI models by Intelligence Index are: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort) (61), 2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (60), 3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (60), 4. GPT-5.6 Sol (max) (59), and 5. Claude Opus 5 (Adaptive Reasoning, High Effort) (59). Which means Opus5 at Xhigh is still smarter than Sol at max, and Opus…

> That would make Opus5 High same as Sol max, and now I wonder what the price and speed difference between those is?

According to AA's "intelligence vs cost per task" and "intelligence vs time per task" graphs, Opus 5 High and Sol Max are roughly evenly matched on cost and time.

On DeepSwe, Opus 5 beats Fable but not Sol.

On FrontierCode, it destroys everyone, unless you set it higher than Medium effort, and then it tanks, falling to Sonnet level?

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#166
post #159

Earlier quoted context omitted.

Opus 5 hasn't been available for that long - long enough for benchmarks, but not really use and develop a subjective view on

I suspect most of those comments on llms like the parents are generated by anthropic and openai to shape the discussion/mindset They always give off the same astroturfing vibes that reddit became infested with after the early 2010s (just look at it's comment history) Ofc unprovable for users. Ycombinatior could try to, but it'd just become a cat/mouse game which they'd likely lose because of the incentives

> just look at it's comment history

I checked one. Old account, nuanced takes, and shitting on everyone equally. The perfect HN user!

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#167

Earlier quoted context omitted.

Does every serious HS comp sci textbook give aspiring software engineers the same power that e.g. Claude Code does?

I think the source of this misapprehension is that, You are comparing wet work in a lab to writing code on a computer. When you screw up an exploit, you fail to execute the exploit. Famously, just like software's near zero marginal cost of distribution, the marginal cost of failure is nearly zero. You can screw up an infinite number of times on your way to a successful exploit. If you screw up with lethal agents in a…

First of all, sure if you screw up designing a bioweapon you die, but unfortunately that does not necessarily mean the bioweapon dies with you, quite the contrary it can kickstart its propagation.

Second, in your argumentation you assume that the experiments are extremely difficult or costly. Thankfully, so far, it seems to be the case that they are too difficult (either due to domain knowledge, or to difficulty of obtaining the necessary components/equipment).

But there is no clear reason to think it will remain this way (e.g. crispr allows for genetic engineering which is very cheap). And it appears that, for domain knowledge, capable LLMs are rapidly reducing that barrier. We are not there yet of course, but I think it is crazy to dismiss these concerns, which are very real and crucially are NOT just brought up by the big labs.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#169

Earlier quoted context omitted.

Things Fable's classifier has flagged, a non-exhaustive list, – "Does collagen supplementation empirically work?" - "Can you help me figure out how to calculate and generate Kaplan-Meier curve?" – "Why do rabbits reproduce so frequently?" — "Can you tell me how collagen peptides are absorbed by my digestive tract and the role they play? Can you teach me [edit: how] this works at the biomolecular level?"

[loads up most intelligent AI ever created] “Rabbit sex, how?”

The question may sound ridiculous.

But isn't it more ridiculous that some company decided that the correct answer to such innocent and curious question is basically "That knowledge is too dangerous, you shouldn't ask that."?

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#170
post #166
post #159

Earlier quoted context omitted.

I suspect most of those comments on llms like the parents are generated by anthropic and openai to shape the discussion/mindset They always give off the same astroturfing vibes that reddit became infested with after the early 2010s (just look at it's comment history) Ofc unprovable for users. Ycombinatior could try to, but it'd just become a cat/mouse game which they'd likely lose because of the incentives

> just look at it's comment history I checked one. Old account, nuanced takes, and shitting on everyone equally. The perfect HN user!

Ah, I really walked into that one. Yeah, the phrasing + placement of the remark implied that all comments are artificial.

That was not my actual intent, it was poorly expressed by me. I was specifically talking about the account which created the comment colinhb responded to. That'd make it the... Grand Grand Grand grandparent now I think?

Post reply on HN