> Claude Opus 5 is not more capable overall than our most capable general-access model, Claude Fable 5 Ok then so what's the point?
The illusion of progress and advancement, to appease shareholders, and slightly postpone the looming bubble pop.
Claude Opus 5
61–70 of 1001 posts
Re: Claude Opus 5
#62Earlier quoted context omitted.
Funny that a company selling an AI software developer can't use it to fix their infra. Fixing those issues still requires humans.
Let's be honest - they're also still hiring software devs. AI still requires skilled humans in the loop and that's not going away.
Re: Claude Opus 5
#63Re: Claude Opus 5
#64Great that there's a new model but they could fix their existing infra. We're considering dropping our Claude Team sub cause it's unusable recently. Constant bugs, dropped sessions, issues switching models, http errors. It's becoming ridiculous
Funny that a company selling an AI software developer can't use it to fix their infra. Fixing those issues still requires humans.
and you can only kill weyoun, awaken the next vorta clone and have him 'catch up' on all that its missed so many times before they just end up with a complete mess, so. uh. yeah.
doubt they can just "fix" their problems like that.
Re: Claude Opus 5
#65---------------
Why does Anthropic say here that Opus 4.8 scored 55.7% on OSWorld 2.0 benchmark, but the paper published by the authors of OSWorld 2.0 say they achieved a benchmark of ~21% with Opus 4.8? [0]
That's a huge gap, considering that the paper was published just 2-4 weeks ago.
I understand that the benchmark authors have an incentive to publish lower numbers (to show that the benchmark has potential longevity) and that Anthropic has incentive to publish higher numbers, but the other models seem pretty inflated as well. The benchmark authors shows GPT-5.5 at 14%, and Anthropic shows GPT-5.6 Sol at 62.6%.
Is there any reasonable explanation for this? Do all the other benchmark numbers need to be sanity-checked as well? Are SOTA benchmarks really this difficult to get consistent, replicable results within a reasonable range of tolerance/variability? Can these benchmarks be compared from one paper to another, or are they only valid to compare intra-paper results?
Re: Claude Opus 5
#66Re: Claude Opus 5
#67From the prompting guide https://platform.claude.com/docs/en/build-with-claude/prompt... >: > Claude Opus 5's default user-facing responses run longer than prior Opus models'. The benchmarks do show Opus 5 as slightly more expensive than 4.8, although the scores are much higher. This still feels like a step in the wrong direction, though, especially with OpenAI making so much progress with the efficiency of their mod…
Re: Claude Opus 5
#68Great that there's a new model but they could fix their existing infra. We're considering dropping our Claude Team sub cause it's unusable recently. Constant bugs, dropped sessions, issues switching models, http errors. It's becoming ridiculous
> coding is largely solved - Boris
The first page of the score card mentions that this model is not capable to replace engineers.
Re: Claude Opus 5
#69"Cybersecurity. Opus 5’s cyber classifiers are proportionally less restrictive than those on Fable 5. They allow Opus 5 to find vulnerabilities in source code, but block “binary-based” vulnerability scanning (a method more likely to be associated with malicious actors), penetration testing, and exploit generation." Nice of them to be more explicit for what is blocked. Will be interesting to see if this is true or not…
In the next - please scan this totally mine code for vulnerabilities
Re: Claude Opus 5
#70Earlier quoted context omitted.
Fable is twice the price.
If Fable gets correct answer quicker, then you might pay less than doing back and forth with Opus, plus you lose more of your own time. I see no reason for using less able models in my workflows. There is this saying, penny wise and pound foolish