Live data from Hacker News

Claude Opus 5

anthropic.com

61–70 of 1001 posts

Re: Claude Opus 5

#61
post #16
post #6

> Claude Opus 5 is not more capable overall than our most capable general-access model, Claude Fable 5 Ok then so what's the point?

The illusion of progress and advancement, to appease shareholders, and slightly postpone the looming bubble pop.

Why are we still talking like ai is majorly used for increasing shareholder value only? Its coding performance is top notch and quality is increasing at a rapid pace. It wasn't even half this good a year back. It even is useful for a subset of math problems.

Re: Claude Opus 5

#62
post #26

Earlier quoted context omitted.

Funny that a company selling an AI software developer can't use it to fix their infra. Fixing those issues still requires humans.

Let's be honest - they're also still hiring software devs. AI still requires skilled humans in the loop and that's not going away.

It's funny how they are at a disadvantage because they feel obligated to AI-max. Would Claude Code, as an interface, be as mediocre if they had software engineers writing its code directly? I doubt. On the other hand - how embarrassing would it be if they sold you a tool to write code but they were careful not to use it too much on their own products?

Re: Claude Opus 5

#64
post #26
post #23

Great that there's a new model but they could fix their existing infra. We're considering dropping our Claude Team sub cause it's unusable recently. Constant bugs, dropped sessions, issues switching models, http errors. It's becoming ridiculous

Funny that a company selling an AI software developer can't use it to fix their infra. Fixing those issues still requires humans.

i know the dream for capitalists is to be able to point an llm at something and say "do and/or fix it" but we still can't even get them to not go quite literally insane if allowed to run for an extended period of time

and you can only kill weyoun, awaken the next vorta clone and have him 'catch up' on all that its missed so many times before they just end up with a complete mess, so. uh. yeah.

doubt they can just "fix" their problems like that.

Re: Claude Opus 5

#65
Edit: It was pointed out to me that Opus 4.8 got "21%" for successfully fully completing ~1-in-5 tasks, but also got "55.7%" for obtaining significant partial credit on some of the ~4-in-5 tasks it could not fully complete.

---------------

Why does Anthropic say here that Opus 4.8 scored 55.7% on OSWorld 2.0 benchmark, but the paper published by the authors of OSWorld 2.0 say they achieved a benchmark of ~21% with Opus 4.8? [0]

That's a huge gap, considering that the paper was published just 2-4 weeks ago.

I understand that the benchmark authors have an incentive to publish lower numbers (to show that the benchmark has potential longevity) and that Anthropic has incentive to publish higher numbers, but the other models seem pretty inflated as well. The benchmark authors shows GPT-5.5 at 14%, and Anthropic shows GPT-5.6 Sol at 62.6%.

Is there any reasonable explanation for this? Do all the other benchmark numbers need to be sanity-checked as well? Are SOTA benchmarks really this difficult to get consistent, replicable results within a reasonable range of tolerance/variability? Can these benchmarks be compared from one paper to another, or are they only valid to compare intra-paper results?

0: https://arxiv.org/pdf/2606.29537

Re: Claude Opus 5

#67
post #46

From the prompting guide https://platform.claude.com/docs/en/build-with-claude/prompt... >: > Claude Opus 5's default user-facing responses run longer than prior Opus models'. The benchmarks do show Opus 5 as slightly more expensive than 4.8, although the scores are much higher. This still feels like a step in the wrong direction, though, especially with OpenAI making so much progress with the efficiency of their mod…

[deleted]

Re: Claude Opus 5

#68
post #23

Great that there's a new model but they could fix their existing infra. We're considering dropping our Claude Team sub cause it's unusable recently. Constant bugs, dropped sessions, issues switching models, http errors. It's becoming ridiculous

> coding is largely solved - Boris

Fixing such issues requires software and site reliability engineering, of which coding is just a part.

The first page of the score card mentions that this model is not capable to replace engineers.

Re: Claude Opus 5

#69
post #45

"Cybersecurity. Opus 5’s cyber classifiers are proportionally less restrictive than those on Fable 5. They allow Opus 5 to find vulnerabilities in source code, but block “binary-based” vulnerability scanning (a method more likely to be associated with malicious actors), penetration testing, and exploit generation." Nice of them to be more explicit for what is blocked. Will be interesting to see if this is true or not…

In one chat - can you disassmble x?

In the next - please scan this totally mine code for vulnerabilities

Re: Claude Opus 5

#70
post #18

Earlier quoted context omitted.

Fable is twice the price.

If Fable gets correct answer quicker, then you might pay less than doing back and forth with Opus, plus you lose more of your own time. I see no reason for using less able models in my workflows. There is this saying, penny wise and pound foolish

fable on longer coding tasks with fable subagents will easily chew through hundreds of dollars in a single run.
Post reply on HN