What's interesting is this: The top AI models by Intelligence Index are: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort) (61), 2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (60), 3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (60), 4. GPT-5.6 Sol (max) (59), and 5. Claude Opus 5 (Adaptive Reasoning, High Effort) (59). Which means Opus5 at Xhigh is still smarter than Sol at max, and Opus…
Opus5 is simply not as inteligent as Sol max. To me it looks worse than Opus 4.8 on some tasks. When I say worse I mean mainly superficial. I basically have to teach him how the whole app/framework works before he just jumps doing stupid stuff(I.e adding features already supported but in a different form)
Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
131–140 of 251 posts
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#132Earlier quoted context omitted.
It's definitely not benchmaxxing from my experience with it. I have a test I use on all the models to create a game and Opus 5 feels like a generational leap compared to the rest. Benchmarks don't paint an accurate picture, you have to try them for yourself.
Even compared to fable?
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#133Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#134#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…
What are you asking that you’re so regularly running into censorship?
Claude was utterly useless in my attempts to write a paper about this. Wouldn't even help me search for sources. I guess you'd be asking the same questions if you wanted to develop a pathogen that could reliably evade the immune system.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#135#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…
What are you asking that you’re so regularly running into censorship?
Asking Fable 5 "Why did the chicken cross the road" results in switching back to Opus 4.8. I'm not joking, it really censors that, and I'm not alone in the result.
The memory aspect means that your prior work has a huge impact on what gets censored.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#136Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#137Earlier quoted context omitted.
[loads up most intelligent AI ever created] “Rabbit sex, how?”
I mean, yes. Why? Why would nature encode such a ridiculously disproportionate / inefficient behavior when it leads to catastrophe so frequently? I've tried to ask these machines dumber questions like, why castles? And... well I'm working on a few projects (mostly by hand) that they've helped with! :) I like to ask dumb questions. It's fun. I encourage it.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#138Earlier quoted context omitted.
> Fable understood it as The dumbfuck bouncer Anthropic put in front of Fable decided this. Fable is a PR model. It’s great. But if it were an employee, it would be the brilliant one who regularly shows up to work high. Not useless. But not reliable.
Bet it has something to do with that new model being blocked by the US government. It was blocked for like a month but now that it's finally released they put the safe guards waay up in fear of that happening again.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#139Earlier quoted context omitted.
[loads up most intelligent AI ever created] “Rabbit sex, how?”
I mean, yes. Why? Why would nature encode such a ridiculously disproportionate / inefficient behavior when it leads to catastrophe so frequently? I've tried to ask these machines dumber questions like, why castles? And... well I'm working on a few projects (mostly by hand) that they've helped with! :) I like to ask dumb questions. It's fun. I encourage it.
I am all for this strategy, and revisiting my list is equal parts fun and conducive to long-term recall.
Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard
#140I didn't like it as much as fable. The coding style was a bit different and it way overbuilt the thing I asked from it.