Live data from Hacker News

Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

claude.com

41–50 of 56 posts

Re: Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

#41

Earlier quoted context omitted.

Agree 100%. This is just another level of obscurity. Security through obscurity... Its annoying, very annoying.

How does security through obscurity apply here?

If you've worked with Codex/CC, you'd have seen it degrade from Fable to Opus, or just not done what you've asked it for.

There are techniques and ways around it. For example, I've found that disabling auto mode in CC sometimes helps (anecdotal, YMMV).

This is all for "security". You can always make it do what you want, it just gets really, really laborious.

Re: Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

#42
I found that Claude / Sol were basically useless when approaching various CTFs that allowed AI tools, but running Qwen and DeepSeek locally and GLM via OpenRouter worked fine. It's blatantly obvious to anyone that even does tangentially cybersecurity related work that Anthropic's position here is stupid and detrimental to security.

Re: Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

#43

The problem is that "cybersecurity" isn't some special task that only your security team does. In the project I maintain, I find bugs and fix bugs. Some of those bugs might result in an LPE. I generate a regression test, then I fix the bug. The problem is that generating a regression test for that type of bug is technically a PoC. I can almost never get Fable to create one. Sometimes Opus 5 punts as well. Same with S…

I’m a web performance engineer — I help clients find and fix site speed issues — but recently I spotted what I thought might be a security/privacy issue. Security not being my specialism, I asked Fable to help me triage and, if necessary, raise the issue with my client.

It refused. It’s so so so adjacent to the work we’d already been doing, but the moment I asked it to help me understand what I thought I’d found, it left me high and dry!

Re: Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

#44

$35M in credits (!) for the Defender Advantage Fund (0xDAF) doesn't sound that much given that e.g. the HAWK attack [1] cost $100k for 1 (albeit very advanced) vulnerability. As a side note, that Golden Eagle wording "bringing a wartime footing to the cyber domain to relentlessly patch vulnerabilities" sounds so AI-written maybe that's what the credits are needed for. (1) https://www.anthropic.com/research/discoverin…

The HAWK attack was on a 3rd round PQC candidate that had been under adversarial review by experts for 3 years. It is an outlier and most vulnerabilities are easier to find.

Re: Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

#45
post #9

> Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders This dance is quite annoying. When Anthropic releases a paid model, the user should control what it does and doesn't do. The other day I had a security incident that required urgent response. ChatGPT and Claude were utterly useless (I quickly attempted to get access to the former's advanced security capabilities but was met with a form I c…

On the other hand, people are using Grok to do things like “nudify” images of random people they find online, including minors.

I’m not a fan of limiting access to models, but the other extreme (no limits or guardrails) is at least as bad if not worse.

Re: Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

#46

Earlier quoted context omitted.

>> In a world where things such as GLM-5.3, DeepSeek-V4-Pro-0813, and Kimi-K3 exists, this is a bit laughable. Is it, though? We don't know what Mythos is really capable of, beyond what Anthropic has told us, and some second-hand accounts from orgs that have been whitelisted. What we do know is that their withholding it from the masses is causing a lot of harm to their reputation and general annoyance. And probably a…

> Is it, though? Yes, it truly is. Open models are extremely capable, as benchmark after benchmark has indicated. Beyond that, for the vast majority of software development (including cybersecurity), the open models are there already. All that without having to pay the hefty Anthropic premium, not to mention all their bullshit with pretending their model is some sort of WMD and their awful uptime (although, to their…

You didn't really address my point. You just... stated your own.

Re: Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

#47
post #40

Earlier quoted context omitted.

> Is it, though? Yes, it truly is. Open models are extremely capable, as benchmark after benchmark has indicated. Beyond that, for the vast majority of software development (including cybersecurity), the open models are there already. All that without having to pay the hefty Anthropic premium, not to mention all their bullshit with pretending their model is some sort of WMD and their awful uptime (although, to their…

As basically every subject matter expert has stated again and again, benchmarks do not tell reveal anything meaningful and have effectively no relation to the model's actual capabilities.

Which is why I actually used all those models.

At work I am stuck with Claude because it is what the employer provides.

In my home setup I use a combination of different models; mostly GLM-5.3, DS-V4-Flash, and MiMo-2.5-pro.

I vastly prefer my home setup.

Re: Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

#48

The problem is that "cybersecurity" isn't some special task that only your security team does. In the project I maintain, I find bugs and fix bugs. Some of those bugs might result in an LPE. I generate a regression test, then I fix the bug. The problem is that generating a regression test for that type of bug is technically a PoC. I can almost never get Fable to create one. Sometimes Opus 5 punts as well. Same with S…

Every one of their models has become absolutely useless, for validation of bugs and the remediations. Unless I’m doing straight forward dev work, I’ve turned to alternate models and harnesses I’ve started building on my own. The frontier models have apparently become so good at security that they can’t be bothered to discuss it with laymen…

Re: Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

#49

Earlier quoted context omitted.

> Is it, though? Yes, it truly is. Open models are extremely capable, as benchmark after benchmark has indicated. Beyond that, for the vast majority of software development (including cybersecurity), the open models are there already. All that without having to pay the hefty Anthropic premium, not to mention all their bullshit with pretending their model is some sort of WMD and their awful uptime (although, to their…

You didn't really address my point. You just... stated your own.

I definitely did. I was talking about how in a world where open models exist and are very much capable, Anthropic is sort of a joke.

You asked if that is true, and speculated about mythos.

Uninterested in speculating about mythos, I explained why it is irrelevant anyway.

This was a successul conversation.

Re: Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders

#50

The problem is that "cybersecurity" isn't some special task that only your security team does. In the project I maintain, I find bugs and fix bugs. Some of those bugs might result in an LPE. I generate a regression test, then I fix the bug. The problem is that generating a regression test for that type of bug is technically a PoC. I can almost never get Fable to create one. Sometimes Opus 5 punts as well. Same with S…

I’m a web performance engineer — I help clients find and fix site speed issues — but recently I spotted what I thought might be a security/privacy issue. Security not being my specialism, I asked Fable to help me triage and, if necessary, raise the issue with my client. It refused. It’s so so so adjacent to the work we’d already been doing, but the moment I asked it to help me understand what I thought I’d found, it…

I run into this several times per week.

Sometimes the context where Fable drops to Opus 4.8 is enough for Opus 4.8 to do a first pass on everything, then I start a new session and Fable will happily refactor/improve it.

Super annoying.

Post reply on HN