Live data from Hacker News

Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

artificialanalysis.ai

31–40 of 251 posts

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#31
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

I'm doing a bunch of x86_64 assembly these days and Fable is simply not allowed to debug it. Hoping Opus 5 has a bit more freedom.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#32
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

I've hit it with intensely benign things; like asking it to make me a web-based client-side word game. I am guessing it saw the dictionary and pattern matched on various words, though ultimately it provided no explanation for why it triggered safeguards.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#33

The more interesting finding is that it's still the second most expensive model (after Fable 5) by a long shot. At least two models (GPT-5.6, Kimi K3) match its score (~1-2% diff) for half the cost.

The chart shows max effort, used mostly by price-insensitive enterprise users. At medium effort it drops to almost half K3’s cost, and is probably sufficient for 95% of coding tasks.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#34
post #22

Earlier quoted context omitted.

What are you asking that you’re so regularly running into censorship?

I'm doing a bunch of x86_64 assembly these days and Fable is simply not allowed to debug it. Hoping Opus 5 has a bit more freedom.

I haven't been using it for long, but so far the refusals seem about on par with how things were on Opus 4.8.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#35
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

Asking fable to read it's own model card triggers this btw. Or asking if mitochondria is the powerhouse of the cell.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#36
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

i do homebrewing and asked it to compare some beer yeasts for me and hit the safeguards because... biology i guess lol

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#37

Earlier quoted context omitted.

5.6 Sol (max) being cheaper than all of these is wild, considering how good the output is too

This must be on API costs, not counting the $100/200 tiers, right?

yes; fyi usage limits on the $200 claude sub correspond to at least $1.2k/week in api tokens

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#38
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

I was profiling a slow machine the other day, and triggered the safeguards.

I've been saying this a lot lately, but it doesn't bites you until it bites you.

The more you use the clanker as a general purpose fix-it tool (goodbye manual NeoVim configuration, you will not be missed!), the more you will find yourself bumping into these safeguards.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#39
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

Just ask math question and it will censor that. Even Misanthrophic employee confirmed that.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#40
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

Working on dimensionl reduction algorithms, I hit it all the time. I'm also trying to port related protocols from single-cell transcriptomics to collective intelligence systems (working with people x reaction matrices as analogous to single-cells cell x gene matrices.

Something between single-cell work and advanced nonlinear DR methods (perhaps used in alignment work?) it always flags me

Post reply on HN