Live data from Hacker News

Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

artificialanalysis.ai

61–70 of 251 posts

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#61
post #35

Earlier quoted context omitted.

Asking fable to read it's own model card triggers this btw. Or asking if mitochondria is the powerhouse of the cell.

Gemini knocks this out of the park, Gemini gang unite. https://share.gemini.google/34vZzlnsmTaL

[dead]

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#62
post #35

Earlier quoted context omitted.

Asking fable to read it's own model card triggers this btw. Or asking if mitochondria is the powerhouse of the cell.

Wait wtf. The mitochondria thing is true. > Why this chat was flagged This model has safety measures that flag specific phrases. This can happen to safe, normal chats. > Your message itself appears to be what’s triggering the safety check. Editing it and retrying may help.

Anything biology-related does this. It even did it when I asked it how eye color works, or something about frogs.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#63
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

I was trying to debug/fix a segfault in the JVM which kept getting flagged

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#64

The more interesting finding is that it's still the second most expensive model (after Fable 5) by a long shot. At least two models (GPT-5.6, Kimi K3) match its score (~1-2% diff) for half the cost.

The chart shows max effort, used mostly by price-insensitive enterprise users. At medium effort it drops to almost half K3’s cost, and is probably sufficient for 95% of coding tasks.

Yeah if there’s one thing that people should really understand it’s that it’s cheaper to have smarter models with less thinking than cheaper models with more thinking.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#65
post #50
post #41

Earlier quoted context omitted.

Every time an online chatter (e.g. "limits are better", "model is better") makes me to reevaluate my principle of never paying Anthropic, I go to the model card, which strengthens my belief in the principle. Why is Anthropic is so hell-bent on this auto/silent downgrade? Do they have a single user who prefers an auto-lobotomization instead of a refusal? Have they learned nothing from the backlash the first time?

Just go to /config. The very second configuration item is “Switch models when a message is flagged” and presumably you want to turn this off. Oh but then you said you never pay Anthropic so you haven’t actually used Claude Code yet. Why would anyone listen to the opinion of a non-user?

> so you haven’t actually used Claude Code yet.

Where do you think the principle came from? I've used claude code for a year, and stopped February this year.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#66

Earlier quoted context omitted.

It shouldn't be surprising OpenAI does have the most compute out of all the major labs. The only reason why Anthropic models are expensive is they are the most in demand models in the world and Anthropic is fighting for compute. The only way to you limit demand for your model is increasing API pricing this is also why Anthropic probably has great margin and probably is profitable compared to OpenAI.

I also suspect there is a price fixing agreement between all of the inference providers for Claude (such as Amazon, Anthropic, Microsoft, etc).

I doubt there's any sort of criminal behavior there - the model is anthropic's up and anthropic probably charges a very expensive license fee that's the same for all of them, and their cogs on compute aren't going to be wildly different, so the main drivers of the cost are roughly the same and they're all offering customers the same end product so the prices would likely also be similar in the end

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#67
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

Fable refuses to work on a login/signup system for example.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#68
post #62

Earlier quoted context omitted.

Wait wtf. The mitochondria thing is true. > Why this chat was flagged This model has safety measures that flag specific phrases. This can happen to safe, normal chats. > Your message itself appears to be what’s triggering the safety check. Editing it and retrying may help.

Anything biology-related does this. It even did it when I asked it how eye color works, or something about frogs.

Probably because the cost of blocking "is mitochondria the powerhouse of the cell" is nearly zero, while the cost of allowing "how do I synthesize the Spanish flu" is approximately infinite.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#69
post #35

Earlier quoted context omitted.

Asking fable to read it's own model card triggers this btw. Or asking if mitochondria is the powerhouse of the cell.

Wait wtf. The mitochondria thing is true. > Why this chat was flagged This model has safety measures that flag specific phrases. This can happen to safe, normal chats. > Your message itself appears to be what’s triggering the safety check. Editing it and retrying may help.

There’s confusion about the classifiers on Fable. They don’t ban chemistry and biology topics they flag as a potential risk, they ban anything related to chemistry or biology at all. This is intentional, and directly stated on the model card, but seems so absurd that there can be assumption it must be a misreading.

Being a researcher somewhat connected to chemistry and biology, Fable has been the most useless model I have ever tried. Essentially all work has instantly downgraded to Opus.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#70
I posted this before but I have a really simple shell tool to keep up with these charts over at

https://github.com/day50-dev/aa-eval-email

This also works

$ curl day50.dev/art-analysis.sh | bash

Artificial analysis knows about my tool and I'm working with them on getting their API improved.

Post reply on HN