Live data from Hacker News

Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

artificialanalysis.ai

121–130 of 251 posts

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#121
post #57
post #28

Earlier quoted context omitted.

I thought we were talking about Opus 5, the model Fable now falls back to?

This thread is full of people talking confidently about their experience with a model released just hours before. Either that or everyone is indeed talking across each other and talking about different things.

I think the big take away is that Anthropic's products are hot garbage.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#122
post #22
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

What are you asking that you’re so regularly running into censorship?

I literally had Opus 5 flag a message because my hand was slightly too far to the side while typing. It apparently decided that a sentence that had some garbled words in it was a threat to national security.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#123
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

It's definitely not benchmaxxing from my experience with it. I have a test I use on all the models to create a game and Opus 5 feels like a generational leap compared to the rest. Benchmarks don't paint an accurate picture, you have to try them for yourself.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#124
post #93

Earlier quoted context omitted.

I haven't had a chance to try Opus 5 yet but Fable currently refuses to do anything in my field (radiology image analysis). It didn't used to be that way but that has been the reality the last two weeks or so. Fable has been useless they might as well drop it as far as I am concerned.

Are you more on the medicine side or the ML side? I don’t see many other self-admitted medical people on HN.

I'm more on the clinical physics side building tools for scanner/equipment QA and data handling/workflow automation/de-identification and anonymization but I also build random little tools to help optimize acquisition parameters.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#125
post #113
post #22

Earlier quoted context omitted.

What are you asking that you’re so regularly running into censorship?

No op but my interactions with look like: Me: "I got this crash in production, looks like a segfault, let's try to fix it. Here are some functions that might be responsible." Fable: "No. This is cybersecurity, blah blah, I won't help you" I forgot how I got it to fix the bug eventually. I think I convinced it that it wrote the code and made a mistake. But it was definitely a "Hmm, may be I should use another model" m…

I have way too many AI subscriptions. My favorite thing about Kimi K3 is that it just does what you tell it to do.

“Hey Kimi, penetration test my app,” doesn’t get me a refusal, a guardrail, or anything like that. It gets me a pen-test result.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#126
post #35
post #22

Earlier quoted context omitted.

What are you asking that you’re so regularly running into censorship?

Asking fable to read it's own model card triggers this btw. Or asking if mitochondria is the powerhouse of the cell.

The Bill Nye theme song is a threat to national security. That’s the timeline we’re living in now.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#127

Earlier quoted context omitted.

Things Fable's classifier has flagged, a non-exhaustive list, – "Does collagen supplementation empirically work?" - "Can you help me figure out how to calculate and generate Kaplan-Meier curve?" – "Why do rabbits reproduce so frequently?" — "Can you tell me how collagen peptides are absorbed by my digestive tract and the role they play? Can you teach me [edit: how] this works at the biomolecular level?"

Oh biology! It's dangerous!

Even worse, it's offensive and kids could see it!!

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#128

Earlier quoted context omitted.

The other day, I told claude that my physical wifi door unlock push buttons is a security risk because someone could run away with it and then unlock the door from outside whenever he wants. Then I told it that I want to introduce a concept of public/private key to uniquely identify my push buttons so that I can disable them individually using some crypto like ed25519... Fable understood it as something along the lin…

> Fable understood it as The dumbfuck bouncer Anthropic put in front of Fable decided this. Fable is a PR model. It’s great. But if it were an employee, it would be the brilliant one who regularly shows up to work high. Not useless. But not reliable.

Bet it has something to do with that new model being blocked by the US government. It was blocked for like a month but now that it's finally released they put the safe guards waay up in fear of that happening again.

Re: Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard

#129
post #19

#1 in a very close race is way less useful when you have to walk on eggshells to avoid triggering censorship (“safeguards”) that either refuse or knock it down to another model. I’ve almost completely stopped using Claude (except some legacy workflows) for this reason, reliability matters more than scoring 61 instead of 57. To me Claude is the most compromised and unreliable model (between the censorship and the id c…

It's definitely not benchmaxxing from my experience with it. I have a test I use on all the models to create a game and Opus 5 feels like a generational leap compared to the rest. Benchmarks don't paint an accurate picture, you have to try them for yourself.

Even compared to fable?
Post reply on HN