Earlier quoted context omitted.
Asking fable to read it's own model card triggers this btw. Or asking if mitochondria is the powerhouse of the cell.
Gemini knocks this out of the park, Gemini gang unite. https://share.gemini.google/34vZzlnsmTaL
Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
61–70 of 251 posts
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#62Earlier quoted context omitted.
Asking fable to read it's own model card triggers this btw. Or asking if mitochondria is the powerhouse of the cell.
Wait wtf. The mitochondria thing is true. > Why this chat was flagged This model has safety measures that flag specific phrases. This can happen to safe, normal chats. > Your message itself appears to be what’s triggering the safety check. Editing it and retrying may help.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#63#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…
What are you asking that you’re so regularly running into censorship?
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#64The more interesting finding is that it's still the second most expensive model (after Fable 5) by a long shot. At least two models (GPT-5.6, Kimi K3) match its score (~1-2% diff) for half the cost.
The chart shows max effort, used mostly by price-insensitive enterprise users. At medium effort it drops to almost half K3’s cost, and is probably sufficient for 95% of coding tasks.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#65Earlier quoted context omitted.
Every time an online chatter (e.g. "limits are better", "model is better") makes me to reevaluate my principle of never paying Anthropic, I go to the model card, which strengthens my belief in the principle. Why is Anthropic is so hell-bent on this auto/silent downgrade? Do they have a single user who prefers an auto-lobotomization instead of a refusal? Have they learned nothing from the backlash the first time?
Just go to /config. The very second configuration item is “Switch models when a message is flagged” and presumably you want to turn this off. Oh but then you said you never pay Anthropic so you haven’t actually used Claude Code yet. Why would anyone listen to the opinion of a non-user?
Where do you think the principle came from? I've used claude code for a year, and stopped February this year.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#66Earlier quoted context omitted.
It shouldn't be surprising OpenAI does have the most compute out of all the major labs. The only reason why Anthropic models are expensive is they are the most in demand models in the world and Anthropic is fighting for compute. The only way to you limit demand for your model is increasing API pricing this is also why Anthropic probably has great margin and probably is profitable compared to OpenAI.
I also suspect there is a price fixing agreement between all of the inference providers for Claude (such as Amazon, Anthropic, Microsoft, etc).
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#67#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…
What are you asking that you’re so regularly running into censorship?
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#68Earlier quoted context omitted.
Wait wtf. The mitochondria thing is true. > Why this chat was flagged This model has safety measures that flag specific phrases. This can happen to safe, normal chats. > Your message itself appears to be what’s triggering the safety check. Editing it and retrying may help.
Anything biology-related does this. It even did it when I asked it how eye color works, or something about frogs.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#69Earlier quoted context omitted.
Asking fable to read it's own model card triggers this btw. Or asking if mitochondria is the powerhouse of the cell.
Wait wtf. The mitochondria thing is true. > Why this chat was flagged This model has safety measures that flag specific phrases. This can happen to safe, normal chats. > Your message itself appears to be what’s triggering the safety check. Editing it and retrying may help.
Being a researcher somewhat connected to chemistry and biology, Fable has been the most useless model I have ever tried. Essentially all work has instantly downgraded to Opus.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#70https://github.com/day50-dev/aa-eval-email
This also works
$ curl day50.dev/art-analysis.sh | bash
Artificial analysis knows about my tool and I'm working with them on getting their API improved.