Live data from Hacker News

Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

artificialanalysis.ai

131–140 of 251 posts

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#131
post #56

What's interesting is this: The top AI models by Intelligence Index are: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort) (61), 2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (60), 3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (60), 4. GPT-5.6 Sol (max) (59), and 5. Claude Opus 5 (Adaptive Reasoning, High Effort) (59). Which means Opus5 at Xhigh is still smarter than Sol at max, and Opus…

Opus5 is simply not as inteligent as Sol max. To me it looks worse than Opus 4.8 on some tasks. When I say worse I mean mainly superficial. I basically have to teach him how the whole app/framework works before he just jumps doing stupid stuff(I.e adding features already supported but in a different form)

Interestingly, if you filter by the coding index, Sol xhigh is the best one, and only Opus5 max is better than Sol max and Sol high.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#132

Earlier quoted context omitted.

It's definitely not benchmaxxing from my experience with it. I have a test I use on all the models to create a game and Opus 5 feels like a generational leap compared to the rest. Benchmarks don't paint an accurate picture, you have to try them for yourself.

Even compared to fable?

Yes. I've since watched a couple review videos on YouTube and all the game tests I've seen Opus 5 produce are incredible even compared to Fable.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#133

Earlier quoted context omitted.

[loads up most intelligent AI ever created] “Rabbit sex, how?”

Rabbits are a good input calorie to output meat ratio, and their excrement makes good cold compost. They breed and litter relatively easily. Theyre also easy to house.

[deleted]

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#134
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

I'm not who you're responding to, but I have a lot of questions about molecular mimicry: evolution pushes pathogens to be shaped like human cell surfaces because that way the immune system won't attack the pathogens (since, by doing so it would also attack the body). It's thought that many autoimmune disorders have an undiscovered pathogen as their cause, one whose mimicry caused such an attack. Discovery of these pathogens could be done computationally, I think. We can catch MHC binding event in process, find the bound protein, figure out which pathogens have genomes that code for proteins of similar shapes (epitopes), and we'd find--I hypothesize--a list of candidate pathogens for the cause of a delayed onset autoimmune disorder. Preventing these infections ahead of time would be a huge win against diseases like multiple sclerosis because without the initial exposure the immune system wouldn't have cloned so many of the cells that are attacking the host.

Claude was utterly useless in my attempts to write a paper about this. Wouldn't even help me search for sources. I guess you'd be asking the same questions if you wanted to develop a pathogen that could reliably evade the immune system.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#135
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

I use lots of biology in my day job.

Asking Fable 5 "Why did the chicken cross the road" results in switching back to Opus 4.8. I'm not joking, it really censors that, and I'm not alone in the result.

The memory aspect means that your prior work has a huge impact on what gets censored.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#136
post #12

Before getting too excited, take a look at the intelligence vs cost matrix: https://artificialanalysis.ai/models?intelligence-index-toke...

Max is lot of extra reasoning. I wonder how many fewer tasks it solves on high. I bet that costs quite a lot less.

[flagged]

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#137

Earlier quoted context omitted.

[loads up most intelligent AI ever created] “Rabbit sex, how?”

I mean, yes. Why? Why would nature encode such a ridiculously disproportionate / inefficient behavior when it leads to catastrophe so frequently? I've tried to ask these machines dumber questions like, why castles? And... well I'm working on a few projects (mostly by hand) that they've helped with! :) I like to ask dumb questions. It's fun. I encourage it.

Lots of species are r-selected, I don't see why that would be considered inefficient. In fact I think there are probably more r-strategists than K-strategists.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#138

Earlier quoted context omitted.

> Fable understood it as The dumbfuck bouncer Anthropic put in front of Fable decided this. Fable is a PR model. It’s great. But if it were an employee, it would be the brilliant one who regularly shows up to work high. Not useless. But not reliable.

Bet it has something to do with that new model being blocked by the US government. It was blocked for like a month but now that it's finally released they put the safe guards waay up in fear of that happening again.

I bet it has something with Dario the drama queen begging the US gov to regulate them(I.e read ask them to put “safety guards” that they already had on hand)

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#139

Earlier quoted context omitted.

[loads up most intelligent AI ever created] “Rabbit sex, how?”

I mean, yes. Why? Why would nature encode such a ridiculously disproportionate / inefficient behavior when it leads to catastrophe so frequently? I've tried to ask these machines dumber questions like, why castles? And... well I'm working on a few projects (mostly by hand) that they've helped with! :) I like to ask dumb questions. It's fun. I encourage it.

Just as stating the wrong thing would be the quickest way to elicit a response in the past, so too can 'dumb' questions prime the context for more complex queries.

I am all for this strategy, and revisiting my list is equal parts fun and conducive to long-term recall.

Post reply on HN