Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
21–30 of 251 posts
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#22#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#23It's new, normal.
"Not fair! They distilled Opus 5!"
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#24#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#25#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…
What are you asking that you’re so regularly running into censorship?
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#26I didn't like it as much as fable. The coding style was a bit different and it way overbuilt the thing I asked from it.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#27#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#28Earlier quoted context omitted.
What are you asking that you’re so regularly running into censorship?
If you even broach language related to biology you’ll get rerouted. I was presenting data in a grid and referred to a grid cell, Fable saw the word “cell” and safeguards kicked in
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#29Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#30Before getting too excited, take a look at the intelligence vs cost matrix: https://artificialanalysis.ai/models?intelligence-index-toke...
Max is lot of extra reasoning. I wonder how many fewer tasks it solves on high. I bet that costs quite a lot less.