Live data from Hacker News

Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

claude.com

31–40 of 56 posts

Re: Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

#31

In a world where things such as GLM-5.3, DeepSeek-V4-Pro-0813, and Kimi-K3 exists, this is a bit laughable. Anthropic needs some model with a fancy name so they can pretend for another while that their model is so powerful it will destroy the world if released. I propose Claude Legend 6.

>In a world where things such as GLM-5.3, DeepSeek-V4-Pro-0813, and Kimi-K3 exists, this is a bit laughable.

This is not just for security, too. Using Fable for anything that could be remotely construed as being connect to chemistry or biology was impossible until a few weeks ago. Now it is slightly better, but still fails on many completely innocuous projects.

So as other models advance, Anthropic's sole frontier offering to entire academic fields remains an Opus that seems to get worse in capability each release. They're starting to become a joke in my field: at a conference a few weeks ago, one presenter laughed when I asked about his use of Fable and pointed out that it would downgrade if the letters 'd', 'n', and 'a' were anywhere near each other, which is not that far from my experience.

Re: Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

#32
$35M in credits (!) for the Defender Advantage Fund (0xDAF) doesn't sound that much given that e.g. the HAWK attack [1] cost $100k for 1 (albeit very advanced) vulnerability.

As a side note, that Golden Eagle wording "bringing a wartime footing to the cyber domain to relentlessly patch vulnerabilities" sounds so AI-written maybe that's what the credits are needed for.

(1) https://www.anthropic.com/research/discovering-cryptographic...

Re: Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

#35

What a bunch of wankery from Anthropic. I already use Sol 5.6 for security auditing and it works great, and as a bonus doesn't give verbal vomit every time.

I suspect this will be seen as a mistep by Anthropic in the long run. Including how their own hyperbole led to the US gov adding export controls.

Re: Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

#39

The problem is that "cybersecurity" isn't some special task that only your security team does. In the project I maintain, I find bugs and fix bugs. Some of those bugs might result in an LPE. I generate a regression test, then I fix the bug. The problem is that generating a regression test for that type of bug is technically a PoC. I can almost never get Fable to create one. Sometimes Opus 5 punts as well. Same with S…

Agree 100%. This is just another level of obscurity. Security through obscurity... Its annoying, very annoying.

How does security through obscurity apply here?

Re: Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

#40

Earlier quoted context omitted.

>> In a world where things such as GLM-5.3, DeepSeek-V4-Pro-0813, and Kimi-K3 exists, this is a bit laughable. Is it, though? We don't know what Mythos is really capable of, beyond what Anthropic has told us, and some second-hand accounts from orgs that have been whitelisted. What we do know is that their withholding it from the masses is causing a lot of harm to their reputation and general annoyance. And probably a…

> Is it, though? Yes, it truly is. Open models are extremely capable, as benchmark after benchmark has indicated. Beyond that, for the vast majority of software development (including cybersecurity), the open models are there already. All that without having to pay the hefty Anthropic premium, not to mention all their bullshit with pretending their model is some sort of WMD and their awful uptime (although, to their…

As basically every subject matter expert has stated again and again, benchmarks do not tell reveal anything meaningful and have effectively no relation to the model's actual capabilities.
Post reply on HN