Live data from Hacker News

Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

artificialanalysis.ai

91–100 of 251 posts

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#91
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

What are you asking that you are NOT regularly running into censorship?

Pretty much everyone I know who uses Claude and works on anything with any level of detail has gotten false-positive flagged

I got flagged for coding in WebAssembly Text, for chrissakes LOL #haX0r

And honestly, Codex handles this better. It says "Things are going to go a little slower because we must perform additional checks on this. Is that OK?" and your only inconvenience is waiting a little longer.

Fable meanwhile just unceremoniously dumps you right into Opus without asking anything, it just tells you "you're in Opus now, sorryyyy!" Lame.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#92
post #35

Earlier quoted context omitted.

Asking fable to read it's own model card triggers this btw. Or asking if mitochondria is the powerhouse of the cell.

Gemini knocks this out of the park, Gemini gang unite. https://share.gemini.google/34vZzlnsmTaL

I would love for Gemini to be competitive but even 3.6 flash doesn’t match sol, or opus 5, or k3

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#93
post #22

Earlier quoted context omitted.

What are you asking that you’re so regularly running into censorship?

I haven't had a chance to try Opus 5 yet but Fable currently refuses to do anything in my field (radiology image analysis). It didn't used to be that way but that has been the reality the last two weeks or so. Fable has been useless they might as well drop it as far as I am concerned.

Are you more on the medicine side or the ML side? I don’t see many other self-admitted medical people on HN.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#94
post #85

Earlier quoted context omitted.

The chart shows max effort, used mostly by price-insensitive enterprise users. At medium effort it drops to almost half K3’s cost, and is probably sufficient for 95% of coding tasks.

For a fair comparison, you should compare to K3 (which AA has not tested yet unfortunately) and GPT 5.6 Sol also on medium or the closest equivalent

I thought k3 medium wasn’t even available yet?

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#96
post #21

I didn't like it as much as fable. The coding style was a bit different and it way overbuilt the thing I asked from it.

It’s crazy that people feel confident making judgments like these when the model’s been out for only a few hours.

It's a pretty easy spot if you're already using an older model daily.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#97
post #14

Very interesting that one of the components is "AA-Omniscience Index" AA-Omniscience Index (higher is better) measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer. This seems to be a good proxy for param size/density and the ranking breaks down as such: Claude Fable 5 (with fallback), Gemini 3.1 Pro Preview, Claude Opus 5 (Ma…

I think grounding has as much if not more go do with it. Google does a great job w/ grounding for obvious reasons

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#98
post #22

Earlier quoted context omitted.

What are you asking that you’re so regularly running into censorship?

Things Fable's classifier has flagged, a non-exhaustive list, – "Does collagen supplementation empirically work?" - "Can you help me figure out how to calculate and generate Kaplan-Meier curve?" – "Why do rabbits reproduce so frequently?" — "Can you tell me how collagen peptides are absorbed by my digestive tract and the role they play? Can you teach me [edit: how] this works at the biomolecular level?"

[loads up most intelligent AI ever created]

“Rabbit sex, how?”

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#99
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

Ask anything related to practical applications for quantum-computing, or space etc. The stuff you can find on Wikipedia.

They were able to solve coding, but not what a real danger is.

Post reply on HN