#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…
What are you asking that you’re so regularly running into censorship?
Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
31–40 of 251 posts
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#32#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…
What are you asking that you’re so regularly running into censorship?
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#33The more interesting finding is that it's still the second most expensive model (after Fable 5) by a long shot. At least two models (GPT-5.6, Kimi K3) match its score (~1-2% diff) for half the cost.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#34Earlier quoted context omitted.
What are you asking that you’re so regularly running into censorship?
I'm doing a bunch of x86_64 assembly these days and Fable is simply not allowed to debug it. Hoping Opus 5 has a bit more freedom.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#35#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…
What are you asking that you’re so regularly running into censorship?
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#36#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…
What are you asking that you’re so regularly running into censorship?
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#37Earlier quoted context omitted.
5.6 Sol (max) being cheaper than all of these is wild, considering how good the output is too
This must be on API costs, not counting the $100/200 tiers, right?
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#38#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…
What are you asking that you’re so regularly running into censorship?
I've been saying this a lot lately, but it doesn't bites you until it bites you.
The more you use the clanker as a general purpose fix-it tool (goodbye manual NeoVim configuration, you will not be missed!), the more you will find yourself bumping into these safeguards.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#39#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…
What are you asking that you’re so regularly running into censorship?
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#40#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…
What are you asking that you’re so regularly running into censorship?
Something between single-cell work and advanced nonlinear DR methods (perhaps used in alignment work?) it always flags me