Live data from Hacker News

Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

artificialanalysis.ai

111–120 of 251 posts

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#111
post #47
post #22

Earlier quoted context omitted.

What are you asking that you’re so regularly running into censorship?

Reverse engineering. Codex sometimes displays an advisory prompt when classifier trips - "Wait longer while we evaluate this request further or use a dumber model". If you do nothing, it'll just take some time and almost always succeed. It does require some brainwashing of the model to get it to the state where model itself agrees to do RE work though. But at least it's all predictable.

I’ve had codex/sol block me due to safeguards tripping without attempting to do anything nefarious. It’s not entirely predicable.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#112
post #109

Earlier quoted context omitted.

> Fable understood it as The dumbfuck bouncer Anthropic put in front of Fable decided this. Fable is a PR model. It’s great. But if it were an employee, it would be the brilliant one who regularly shows up to work high. Not useless. But not reliable.

> Fable is a PR model. Yeah, Fable is Anthropic's Cybertruck.

Not quite. Fable is a Model S. The problem is you have to buy a Cybertruck to get it.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#113
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

No op but my interactions with look like:

Me: "I got this crash in production, looks like a segfault, let's try to fix it. Here are some functions that might be responsible."

Fable: "No. This is cybersecurity, blah blah, I won't help you"

I forgot how I got it to fix the bug eventually. I think I convinced it that it wrote the code and made a mistake. But it was definitely a "Hmm, may be I should use another model" moment".

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#114
Honestly, who the fuck cares? These leaderboards are meaningless for brand new models. If we were looking at longitudinal data collected over the course of a year or even a quarter or month, this would have some value. Brand new model from established provider shoots to top of charts? This means nothing more than an already famous band briefly topping the charts with their latest song.

It baffles me that intelligent people deploying AI think momentary popularity is a meaningful signal. It's just twitch reactions * FOMO.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#115
post #93

Earlier quoted context omitted.

I haven't had a chance to try Opus 5 yet but Fable currently refuses to do anything in my field (radiology image analysis). It didn't used to be that way but that has been the reality the last two weeks or so. Fable has been useless they might as well drop it as far as I am concerned.

Are you more on the medicine side or the ML side? I don’t see many other self-admitted medical people on HN.

I've noticed a few self identify and other random occupations, cool to see those fields checking out this tech at a deeper surface level.

I'm just a dev, but I appreciate the insights from other professionals.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#116
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

I told Claude Opus 5.0 to use a global api key for a PFAAS to deploy some web applications in a test environment and import some data into them. It balked at using a global api key because the security issues surrounding the permissiveness run afoul of it's sensibilities.

I have done this task with Opus 4.5, Opus 4.6, Opus 4.7, Opus 4.8 and Fable, without issues.

I have done this task with Codex 5.4, Codex 5.5, and Sol 5.6 without issues.

Opus 5 is too cautious to be productive for me. It needs more tuning.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#117
post #56

What's interesting is this: The top AI models by Intelligence Index are: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort) (61), 2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (60), 3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (60), 4. GPT-5.6 Sol (max) (59), and 5. Claude Opus 5 (Adaptive Reasoning, High Effort) (59). Which means Opus5 at Xhigh is still smarter than Sol at max, and Opus…

Opus5 is simply not as inteligent as Sol max. To me it looks worse than Opus 4.8 on some tasks. When I say worse I mean mainly superficial. I basically have to teach him how the whole app/framework works before he just jumps doing stupid stuff(I.e adding features already supported but in a different form)

"him"? Have we reached that dystopia level?

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#118
post #21

I didn't like it as much as fable. The coding style was a bit different and it way overbuilt the thing I asked from it.

It’s crazy that people feel confident making judgments like these when the model’s been out for only a few hours.

Its even crazier that people are sitting here trying to calculate intelligence per dollar from metrics. At least first impressions have more basis in real performance.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#119
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

The utter meme-think direction this company takes with regards to sycophancy of its models is disgusting.
Post reply on HN