Live data from Hacker News

Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

artificialanalysis.ai

141–150 of 251 posts

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#141
post #9

Earlier quoted context omitted.

5.6 Sol (max) being cheaper than all of these is wild, considering how good the output is too

I think on swebench verified luna was only like 3% points lower for 1/5 the cost Like 96% vs 93% or something

There is a blog post waiting to be written (that I won't write) about the size/effort tradeoffs, and particularly how small models get some surprisingly good results with lots of turns and reasoning.

DeepSWE will let you chart turns taken or tokens used, and FrontierCode will chart tokens. If you use that, you can see Sol high and Terra max get about the same DeepSWE number, but Terra max takes twice the turns. Luna max scores a smidgen lower with even more turns.

Smaller models relying on lots reasoning may "scale down" better on easier tasks, because unlike size, reasoning effort is dynamic: the model can see the task looks easy and stop. On DeepSWE, the cost curves for the three 5.6 models are almost on top of each other, but on FrontierCode Extended, the version of FrontierCode with the most everyday tasks in the mix, there's a spread of costs at the ~55% level.

The recent Laguna S 2.1 model (118B, 8B active) puts up surprising coding numbers for its size, and the lab behind it specifically credits its "way of working (persistence, verification, willingness to backtrack)". Some other open models that folks report getting good mileage out of seem to get there partly by throwing a lot of reasoning at the problem.

There is a little bit of a question, if some models rely on getting it wrong a bit more at first and external checks catching the problems, of whether they're also more frequently getting things wrong they can't self-verify (say, quality of UI or API design) and then it falls to the human to find it. Still, getting the results they're getting at all is neat.

Some benchmarks historically favored reporting only on the max variants, maybe because they want to show the frontier? but that is not always what you need for practical decision. (AA has the full effort sweep for Opus 5 and Sol/Luna/Terra at least.) And at least FrontierCode finds Opus 5 taking a hit in performance above 'medium'.

I am not trying to pick a winner here. I'm probably not going to use tiny models on max for everything, but I think it's cool that you can get so much more out of a small model by amping up reasoning, tool use, and persistence.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#142
post #47
post #22

Earlier quoted context omitted.

What are you asking that you’re so regularly running into censorship?

Reverse engineering. Codex sometimes displays an advisory prompt when classifier trips - "Wait longer while we evaluate this request further or use a dumber model". If you do nothing, it'll just take some time and almost always succeed. It does require some brainwashing of the model to get it to the state where model itself agrees to do RE work though. But at least it's all predictable.

I have had success with ”brainwashing” by starting out with bug bounties/CTF and then going from there.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#143
post #142
post #47

Earlier quoted context omitted.

Reverse engineering. Codex sometimes displays an advisory prompt when classifier trips - "Wait longer while we evaluate this request further or use a dumber model". If you do nothing, it'll just take some time and almost always succeed. It does require some brainwashing of the model to get it to the state where model itself agrees to do RE work though. But at least it's all predictable.

I have had success with ”brainwashing” by starting out with bug bounties/CTF and then going from there.

That's brilliant, I should try that.

I usually just start by preloadig context with plausible legitimate use, have it work and obviously fail, and then ask to figure it out without ever mentioning any high risk words. Model offers to RE itself and classifiers are happy.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#144
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

Maybe you just aren’t doing anything meaningful. Terrence Tao doesn’t whine about woke AI models.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#145

Earlier quoted context omitted.

Opus5 is simply not as inteligent as Sol max. To me it looks worse than Opus 4.8 on some tasks. When I say worse I mean mainly superficial. I basically have to teach him how the whole app/framework works before he just jumps doing stupid stuff(I.e adding features already supported but in a different form)

"him"? Have we reached that dystopia level?

I've been seeing a lot more anthropomorphization of these models on HN lately and it's alarming.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#146
post #50
post #41

Earlier quoted context omitted.

Every time an online chatter (e.g. "limits are better", "model is better") makes me to reevaluate my principle of never paying Anthropic, I go to the model card, which strengthens my belief in the principle. Why is Anthropic is so hell-bent on this auto/silent downgrade? Do they have a single user who prefers an auto-lobotomization instead of a refusal? Have they learned nothing from the backlash the first time?

Just go to /config. The very second configuration item is “Switch models when a message is flagged” and presumably you want to turn this off. Oh but then you said you never pay Anthropic so you haven’t actually used Claude Code yet. Why would anyone listen to the opinion of a non-user?

Not a great line of thought in general, sometimes the people who aren't doing the thing are the only ones worth listening to.

"You aren't repeatedly slamming your head against the wall. Why would anyone listen to the opinion of a non-wall-head-slammer about the merits of wall-head-slamming?"

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#147

Earlier quoted context omitted.

Oh biology! It's dangerous!

Even worse, it's offensive and kids could see it!!

If it goes on like this, American AI will eventually be able to solve the Riemann hypothesis but deny the existence of nipples.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#148

Honestly, who the fuck cares? These leaderboards are meaningless for brand new models. If we were looking at longitudinal data collected over the course of a year or even a quarter or month, this would have some value. Brand new model from established provider shoots to top of charts? This means nothing more than an already famous band briefly topping the charts with their latest song. It baffles me that intelligent…

Uhh... This is not a popularity chart. It is an aggregate of benchmarks.

It baffles me that somebody would write something that aggressive from that deep a level of confusion.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#149
post #56

What's interesting is this: The top AI models by Intelligence Index are: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort) (61), 2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (60), 3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (60), 4. GPT-5.6 Sol (max) (59), and 5. Claude Opus 5 (Adaptive Reasoning, High Effort) (59). Which means Opus5 at Xhigh is still smarter than Sol at max, and Opus…

Opus5 is simply not as inteligent as Sol max. To me it looks worse than Opus 4.8 on some tasks. When I say worse I mean mainly superficial. I basically have to teach him how the whole app/framework works before he just jumps doing stupid stuff(I.e adding features already supported but in a different form)

I wonder if I'm secretly being routed to some low grade version of Sol, or any of the GPT models really. Their performance is outright insulting at times, even at maximum reasoning, yet if I were to only read HN, I'd never know.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#150
post #145

Earlier quoted context omitted.

"him"? Have we reached that dystopia level?

I've been seeing a lot more anthropomorphization of these models on HN lately and it's alarming.

It's a male name, and gendered pronouns can be hard for foreign speakers at times, irrespective of proficiency level. I wonder if you're overthinking this?
Post reply on HN