Live data from Hacker News

Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

artificialanalysis.ai

51–60 of 251 posts

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#51
post #22

Earlier quoted context omitted.

What are you asking that you’re so regularly running into censorship?

Literally anything related to nutrition, athletic performance, etc especially if you ask it for research or sources

Writing an implementation of a board game and one of the cards is called "microbes". Instantly knocked down to a lower tier model whenever it encounters that keyword because clearly bioweapons. Sigh.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#52
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

The other day, I told claude that my physical wifi door unlock push buttons is a security risk because someone could run away with it and then unlock the door from outside whenever he wants. Then I told it that I want to introduce a concept of public/private key to uniquely identify my push buttons so that I can disable them individually using some crypto like ed25519...

Fable understood it as something along the lines of:

"introducing" "security risk" "using software" to "unlock door" YOU ARE FLAGGED

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#53
post #35
post #22

Earlier quoted context omitted.

What are you asking that you’re so regularly running into censorship?

Asking fable to read it's own model card triggers this btw. Or asking if mitochondria is the powerhouse of the cell.

Wait wtf. The mitochondria thing is true.

> Why this chat was flagged This model has safety measures that flag specific phrases. This can happen to safe, normal chats.

> Your message itself appears to be what’s triggering the safety check. Editing it and retrying may help.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#54
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

I routinely get into blocks when running medicine-related material through it.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#55
post #35
post #22

Earlier quoted context omitted.

What are you asking that you’re so regularly running into censorship?

Asking fable to read it's own model card triggers this btw. Or asking if mitochondria is the powerhouse of the cell.

Gemini knocks this out of the park, Gemini gang unite.

https://share.gemini.google/34vZzlnsmTaL

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#56
What's interesting is this:

The top AI models by Intelligence Index are: 1. Claude Opus 5 (Adaptive Reasoning, Max Effort) (61), 2. Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) (60), 3. Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) (60), 4. GPT-5.6 Sol (max) (59), and 5. Claude Opus 5 (Adaptive Reasoning, High Effort) (59).

Which means Opus5 at Xhigh is still smarter than Sol at max, and Opus5 at High is equal to Sol at max.

That would make Opus5 High same as Sol max, and now I wonder what the price and speed difference between those is?

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#57
post #28

Earlier quoted context omitted.

If you even broach language related to biology you’ll get rerouted. I was presenting data in a grid and referred to a grid cell, Fable saw the word “cell” and safeguards kicked in

I thought we were talking about Opus 5, the model Fable now falls back to?

This thread is full of people talking confidently about their experience with a model released just hours before.

Either that or everyone is indeed talking across each other and talking about different things.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#58
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

I was flagged for basically answering yes to what Claude suggested to do, which was test commands on a port for my code for tests we had been discussing. I really think it was flagged simply because the words test and port were in the prompt. Their filters are pathetically poor.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#59
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

I haven't had a chance to try Opus 5 yet but Fable currently refuses to do anything in my field (radiology image analysis). It didn't used to be that way but that has been the reality the last two weeks or so. Fable has been useless they might as well drop it as far as I am concerned.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#60
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

Trying to have it do some rework on a patch to Postgres I'm working on, it just completely shuts down. The reported issues were with privileged escalation and I was instructing it on how to fix.
Post reply on HN