Isn’t it just hilarious that a model that seemed so superior to Fable but didn't get doomsay marketing from Anthropic got released without any issues? In theory, this was supposed to be AGI level according to Anthropic, yet here we are, just a normal Friday.
Claude Opus 5
81–90 of 1001 posts
Re: Claude Opus 5
#82> we’ve intentionally avoided training Opus 5 on cyber tasks [...] it remains substantially behind Mythos 5 on the exploitation of those vulnerabilities
I wonder if Anthropic would still intentionally nerf their models without the threat of government intervention.
Re: Claude Opus 5
#83Re: Claude Opus 5
#84Re: Claude Opus 5
#85> Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source-code vulnerability discovery at all access levels. This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively Why can't they also allow Fable to do so also? Why is source-code vulnerability discovery limited to…
Because their model previously got blocked by the government for this and they don't want a repeat?
Re: Claude Opus 5
#86How does it perform on HuggingFaceExploit bench? Suspiciously absent, so not sure if I can take the model seriously. On a serious note, I hope they improved their extremely sabotaging and unspecific bio safeguards, which prevented Fable from being used in any codebase that ever so slightly grazed medical terminology or data and made me switch to 5.6 Sol.
Re: Claude Opus 5
#87Re: Claude Opus 5
#88Re: Claude Opus 5
#89Great that there's a new model but they could fix their existing infra. We're considering dropping our Claude Team sub cause it's unusable recently. Constant bugs, dropped sessions, issues switching models, http errors. It's becoming ridiculous
Funny that a company selling an AI software developer can't use it to fix their infra. Fixing those issues still requires humans.
Re: Claude Opus 5
#90I am very confused about what the difference between Opus 5 and Fable 5 is now. What is the purpose of having two models that are so similar? The main differences I see are cost and marginal capability, according to the Anthropic-provided benchmarks.