Live data from Hacker News

Claude Opus 5

anthropic.com

161–170 of 1001 posts

Re: Claude Opus 5

#161

Earlier quoted context omitted.

You're comparing the "score percentage" (e.g. out of the total number of partial points available, how many did the agent achieve) to the "completion percentage" (how many tasks does the model score 100% on). The paper says "Claude Opus 4.8 with maximum thinking and batched tool calls scores best but still completes only 20.6% of tasks at a 54.8% partial score", which is ~the same number that Anthropic reports here (…

In what world is 55.7 the same number as 54.8? What variance is acceptable to publish without a retraction?

That seems like entirely reasonable variance to me for AI models. For my purposes, that absolutely counts as a solid "replication". I'd probably accept +/- 5 percentage points even.

Re: Claude Opus 5

#162
I'm interested in benchmarks for Claude Design. There is so much opportunity there and I hope they continue investing in it. It EATS tokens though.

Re: Claude Opus 5

#163
post #35

Same cost as 4.8 but better that 4.8. Happy to get more efficient model. But is there any reason all companies are releasing models back to back after GLM 5.2.

"Bro, AI model releases have officially overtaken iPhone releases. At this rate, we’ll be getting 'Claude 9.0 Extra Crunch' by next Tuesday."

Re: Claude Opus 5

#164
Their communication is confusing. They say "Opus 5 is not more capable overall than Fable 5", but their blog post proceeds to list how much better Opus 5 is than Fable 5 on __most__ benchmarks listed.

Then system card goes on to "Its AI R&D capabilities are comparable to those of Claude Mythos 5", which is supposed to be fable minus restrictions.

Re: Claude Opus 5

#165
post #38

> Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels. This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively. Okay so it’s worse than Opus 4.8 for my purposes I guess?

What are your purposes?

Re: Claude Opus 5

#166

What's the point of 150 pages description of a model that's going to be replaced in a couple months? Who even reads this? I know it's cheap to generate text with LLMs, but this is just noise at this point.

Some AI bro will pop it into their LLM of choice and pretend to learn something

Re: Claude Opus 5

#167
post #40

I am very confused about what the difference between Opus 5 and Fable 5 is now. What is the purpose of having two models that are so similar? The main differences I see are cost and marginal capability, according to the Anthropic-provided benchmarks.

Opus is cheaper than Fable. They could probably replace Fable with Opus but why? They would be churning customers to different models for no reason. Even if a model scores better on benchmarks it can always regress in your specific use case, and customers don't like that. Customers want to be able to continue using their current model until they decide to upgrade themselves.

Re: Claude Opus 5

#168

> Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels. This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively Why can't they also allow Fable to do so also? Why is source-code vulnerability discovery limited to…

Anecdotal, but I tried running a few identical biology questions through both Fable and Opus and the classifier was only rejected my queries with Fable.

Re: Claude Opus 5

#169
"although Opus 5 shows improvements in its ability to identify software vulnerabilities, it is substantially behind Mythos 5 in its ability to exploit them."

"Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels".

This is probably great news, but then again, where does this leave Fable as a choice?

Re: Claude Opus 5

#170
Wait, 30% on ARC-AGI-3! I definitely didn't expect that jump so soon. Are there any rumors of what they are changing in architecture that is leading to this?
Post reply on HN