Live data from Hacker News

Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

artificialanalysis.ai

71–80 of 251 posts

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#71
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

The better question is how would you not? I got demoted to opus from fable for asking if a cancer vaccine I saw on YouTube based on frog bacteria was a real thing. I've gotten it for asking how encryption works. It's incredibly touchy.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#72
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

I took a photo of a rose bush and asked “what’s going on with this rose bush” which triggered a downgrade to Opus. It diagnosed it with rose rosette disease.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#73
post #62

Earlier quoted context omitted.

Anything biology-related does this. It even did it when I asked it how eye color works, or something about frogs.

Probably because the cost of blocking "is mitochondria the powerhouse of the cell" is nearly zero, while the cost of allowing "how do I synthesize the Spanish flu" is approximately infinite.

    >  while the cost of allowing "how do I synthesize the Spanish flu" is approximately infinite
I've heard this sentiment repeated elsewhere, but why? What makes you think that's the case?

Under this rationale, every serious HS textbook has "approximately infinite" risk. That's clearly not so. Why is this any special?

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#74
post #44

Earlier quoted context omitted.

I also suspect there is a price fixing agreement between all of the inference providers for Claude (such as Amazon, Anthropic, Microsoft, etc).

"Price fixing" isn't the correct term here but yes, it's very common to have the same price across different retailers/resellers.

There is a difference between the market discovering a price and a bunch of retailers/resellers entering an agreement to sell at a specific price.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#75
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

Things Fable's classifier has flagged, a non-exhaustive list,

    – "Does collagen supplementation empirically work?"
    - "Can you help me figure out how to calculate and generate Kaplan-Meier curve?"
    – "Why do rabbits reproduce so frequently?"
    — "Can you tell me how collagen peptides are absorbed by my digestive tract and the role they play? Can you teach me [edit: how] this works at the biomolecular level?"

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#76
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

I had a long session about SQL with Fable and at some point, it started to falsely trigger censorship for any message I write in that conversation, even for the simple string: "random message."

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#77

Earlier quoted context omitted.

Probably because the cost of blocking "is mitochondria the powerhouse of the cell" is nearly zero, while the cost of allowing "how do I synthesize the Spanish flu" is approximately infinite.

> while the cost of allowing "how do I synthesize the Spanish flu" is approximately infinite I've heard this sentiment repeated elsewhere, but why? What makes you think that's the case? Under this rationale, every serious HS textbook has "approximately infinite" risk. That's clearly not so. Why is this any special?

Does every serious HS comp sci textbook give aspiring software engineers the same power that e.g. Claude Code does?

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#78

Earlier quoted context omitted.

I also suspect there is a price fixing agreement between all of the inference providers for Claude (such as Amazon, Anthropic, Microsoft, etc).

I doubt there's any sort of criminal behavior there - the model is anthropic's up and anthropic probably charges a very expensive license fee that's the same for all of them, and their cogs on compute aren't going to be wildly different, so the main drivers of the cost are roughly the same and they're all offering customers the same end product so the prices would likely also be similar in the end

Do you really think there is nothing someone could do to make it a fraction of a percentage cheaper to serve like having access to cheaper electricity or a more mature cloud management software. Even saving a fraction of a penny on the prices can make a different due to how much volume people are paying for.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#79
post #56

What's interesting is this: The top AI models by Intelligence Index are: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort) (61), 2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (60), 3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (60), 4. GPT-5.6 Sol (max) (59), and 5. Claude Opus 5 (Adaptive Reasoning, High Effort) (59). Which means Opus5 at Xhigh is still smarter than Sol at max, and Opus…

Opus5 is simply not as inteligent as Sol max. To me it looks worse than Opus 4.8 on some tasks. When I say worse I mean mainly superficial. I basically have to teach him how the whole app/framework works before he just jumps doing stupid stuff(I.e adding features already supported but in a different form)

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#80
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

I wanted to explore some battery chemistry with Fable. It decided I was a terrorist, and then blew my session’s usage credits telling me to fuck off.

The current state of guardrails seems to be entirely about marketing to investors at the cost of customers. I’m switching to open models when my subscription expires. Almost everyone I know, including those with access to Mythos, plan the same at the earliest opportunity. (Or until one of the SOTA models leaks.)

Post reply on HN