Live data from Hacker News

Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

artificialanalysis.ai

151–160 of 251 posts

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#151
post #145

Earlier quoted context omitted.

"him"? Have we reached that dystopia level?

I've been seeing a lot more anthropomorphization of these models on HN lately and it's alarming.

Not all HN visitors are native English speakers and in some languages "it" doesn't construct well with verbs, thus thought frameworks forms through usage of him/her. Nothing more to see I suppose.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#152

Earlier quoted context omitted.

Opus5 is simply not as inteligent as Sol max. To me it looks worse than Opus 4.8 on some tasks. When I say worse I mean mainly superficial. I basically have to teach him how the whole app/framework works before he just jumps doing stupid stuff(I.e adding features already supported but in a different form)

I wonder if I'm secretly being routed to some low grade version of Sol, or any of the GPT models really. Their performance is outright insulting at times, even at maximum reasoning, yet if I were to only read HN, I'd never know.

I've mostly actually stuck to low reasoning for most tasks since it seems to do a surprisingly good job even at low for the stuff I've been throwing at it, and I literally switched directly from Fable 5 to Sol more or less.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#153
post #56

What's interesting is this: The top AI models by Intelligence Index are: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort) (61), 2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (60), 3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (60), 4. GPT-5.6 Sol (max) (59), and 5. Claude Opus 5 (Adaptive Reasoning, High Effort) (59). Which means Opus5 at Xhigh is still smarter than Sol at max, and Opus…

Opus5 is simply not as inteligent as Sol max. To me it looks worse than Opus 4.8 on some tasks. When I say worse I mean mainly superficial. I basically have to teach him how the whole app/framework works before he just jumps doing stupid stuff(I.e adding features already supported but in a different form)

Sol is a complete mess for me.

It only works on end to end tasks in fresh codebases.

Otherwise it cannot follow instructions, changes and deletes unrelated features or does sloppy work to mark a task completed while leaving a compromised codebase.

I could not get Sol to finish a feature in a complex code base without several loops of fixing and reverting

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#154

Earlier quoted context omitted.

Opus5 is simply not as inteligent as Sol max. To me it looks worse than Opus 4.8 on some tasks. When I say worse I mean mainly superficial. I basically have to teach him how the whole app/framework works before he just jumps doing stupid stuff(I.e adding features already supported but in a different form)

"him"? Have we reached that dystopia level?

i prefer to call her Claudia

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#155

Honestly, who the fuck cares? These leaderboards are meaningless for brand new models. If we were looking at longitudinal data collected over the course of a year or even a quarter or month, this would have some value. Brand new model from established provider shoots to top of charts? This means nothing more than an already famous band briefly topping the charts with their latest song. It baffles me that intelligent…

[deleted]

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#156
post #145

Earlier quoted context omitted.

"him"? Have we reached that dystopia level?

I've been seeing a lot more anthropomorphization of these models on HN lately and it's alarming.

Every french person I've talked to IRL, for example, calls Claude "him." It's partly a language thing.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#157
post #56

What's interesting is this: The top AI models by Intelligence Index are: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort) (61), 2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (60), 3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (60), 4. GPT-5.6 Sol (max) (59), and 5. Claude Opus 5 (Adaptive Reasoning, High Effort) (59). Which means Opus5 at Xhigh is still smarter than Sol at max, and Opus…

Opus5 is simply not as inteligent as Sol max. To me it looks worse than Opus 4.8 on some tasks. When I say worse I mean mainly superficial. I basically have to teach him how the whole app/framework works before he just jumps doing stupid stuff(I.e adding features already supported but in a different form)

Opus 5 hasn't been available for that long - long enough for benchmarks, but not really use and develop a subjective view on

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#158
post #153

Earlier quoted context omitted.

Opus5 is simply not as inteligent as Sol max. To me it looks worse than Opus 4.8 on some tasks. When I say worse I mean mainly superficial. I basically have to teach him how the whole app/framework works before he just jumps doing stupid stuff(I.e adding features already supported but in a different form)

Sol is a complete mess for me. It only works on end to end tasks in fresh codebases. Otherwise it cannot follow instructions, changes and deletes unrelated features or does sloppy work to mark a task completed while leaving a compromised codebase. I could not get Sol to finish a feature in a complex code base without several loops of fixing and reverting

Might be an indication that your task unit is too unstructured or your code base is a mess.

Is your actual code doing something complex or is this incidental complexity?

Relying on the model’s “intelligence” to patch over these issues hasn’t proven to be a reliable strategy for me. Of course, this might not apply to you, just my 2 paisa.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#159

Earlier quoted context omitted.

Opus5 is simply not as inteligent as Sol max. To me it looks worse than Opus 4.8 on some tasks. When I say worse I mean mainly superficial. I basically have to teach him how the whole app/framework works before he just jumps doing stupid stuff(I.e adding features already supported but in a different form)

Opus 5 hasn't been available for that long - long enough for benchmarks, but not really use and develop a subjective view on

I suspect most of those comments on llms like the parents are generated by anthropic and openai to shape the discussion/mindset

They always give off the same astroturfing vibes that reddit became infested with after the early 2010s (just look at it's comment history)

Ofc unprovable for users. Ycombinatior could try to, but it'd just become a cat/mouse game which they'd likely lose because of the incentives

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#160

Earlier quoted context omitted.

I also suspect there is a price fixing agreement between all of the inference providers for Claude (such as Amazon, Anthropic, Microsoft, etc).

I doubt there's any sort of criminal behavior there - the model is anthropic's up and anthropic probably charges a very expensive license fee that's the same for all of them, and their cogs on compute aren't going to be wildly different, so the main drivers of the cost are roughly the same and they're all offering customers the same end product so the prices would likely also be similar in the end

Requiring the exact same product to be set at a specific price across providers would not be criminal behavior, lol
Post reply on HN