Live data from Hacker News

Evaluation of Claude Mythos Preview's cyber capabilities

aisi.gov.uk

21–30 of 35 posts

Re: Evaluation of Claude Mythos Preview's cyber capabilities

#21

Earlier quoted context omitted.

> Uh, so those charts don’t look… particularly impressive at all to anyone else? I suspect Anthropic gave them early access hoping for a marketing win and ended up with their arse being served to them on a plate. All rather predictable really. As you say "more compute needed" as the default answer from the AI companies is completely unsustainable. As for the value of Anthropic blog posts, well...

The CTF charts are the less interesting result. (article: "Even expert-level CTFs only test specific skills in isolation.") Models converging at non-expert level isn't a knock on Mythos, it's the benchmark saturating. Of course GPT-5 matches it there. The actual result is TLO, and "only 6 more steps" in OP misreads how sequential attack chains work. These aren't independent puzzles. Each step gates the next. Averagin…

> Compute costs fall reliably, so what matters is the capability at a given price point in 18 months, not today.

The underlying point still stands, namely that "more compute" as the default answer is not sustainable.

Why ?

Because even if we accept the unlikely dream that GPU prices will magically take a nose-dive, you still need somewhere to put all those servers stuffed with GPUs.

That means datacentres.

And "more datacentres" is absolutely not sustainable.

The cooling needs, the power needs, the land needs..... none of it is remotely sustainable.

Re: Evaluation of Claude Mythos Preview's cyber capabilities

#22
This is great data that shows a more realistic view of what Mythos is capable of and exactly what I’ve been hoping to see on the tail of the hype-train.

These details are what are actually important to defenders like myself.

As others have pointed out the limitations are revealing but also the fact that it even made it to the end (despite the cost) is impressive.

I’m hoping that we can get to a point in the future where any skepticism around claims made by the companies producing these models isn’t met with immediate downvotes and accusations of being a Luddite.

Re: Evaluation of Claude Mythos Preview's cyber capabilities

#23
post #14

Once again an evaluation missing confidence intervals. “continued improvement” and “significant improvement” but without any significance testing is moot. With many colleagues (including from AISI themselves!), we recently reviewed 445 the AI benchmarks & evaluations from the past few years. Our work was published at NeurIPS ( https://openreview.net/pdf?id=mdA5lVvNcU ) and we made eight recommendations for better eva…

YES PLEASE! AI's rigor in terms of evaluation of ML systems has been only barely improving in the past 15 years.

Thanks for the concrete recommendations; unfortunately, most of these will fall flat, because nobody teaches how to do these, why they are important, etc.

Re: Evaluation of Claude Mythos Preview's cyber capabilities

#24

Earlier quoted context omitted.

The CTF charts are the less interesting result. (article: "Even expert-level CTFs only test specific skills in isolation.") Models converging at non-expert level isn't a knock on Mythos, it's the benchmark saturating. Of course GPT-5 matches it there. The actual result is TLO, and "only 6 more steps" in OP misreads how sequential attack chains work. These aren't independent puzzles. Each step gates the next. Averagin…

> Compute costs fall reliably, so what matters is the capability at a given price point in 18 months, not today. The underlying point still stands, namely that "more compute" as the default answer is not sustainable. Why ? Because even if we accept the unlikely dream that GPU prices will magically take a nose-dive, you still need somewhere to put all those servers stuffed with GPUs. That means datacentres. And "more…

"datacenters are unsustainable" is a different conversation entirely.

The premise that inference compute must scale linearly with capability isn't supported by what's actually happening. Distillation, quantization, and architectural efficiency gains routinely let you run yesterday's frontier capability at a fraction of the cost and hardware. GPT-3.5 level performance runs on a phone now. The 100M token budget Mythos used here will not require 100M tokens worth of 2026 hardware forever.

On datacenters specifically: yes, they require power, cooling, and land. So does every other piece of industrial infrastructure humanity has ever built and then incrementally made more efficient. The energy per FLOP has been dropping for decades and continues to. You can argue the rate of buildout is concerning, and that's a reasonable discussion. But "not sustainable" as a flat declaration requires you to believe efficiency gains will stall, energy production won't expand, and cooling technology will stay static, all simultaneously. That's a much stronger claim than it sounds.

None of which has anything to do with whether the AISI results are significant. They are.

Re: Evaluation of Claude Mythos Preview's cyber capabilities

#25

Earlier quoted context omitted.

This article reinforces something I've heard a lot of people say for a while now and what I've personally felt. Claude and GPT are fairly evenly matched on any individual task (GPT might even be a little better), but Claude is far more autonomous. So with that said, I think the graph under the "Cyber range results" is the important one. The ones at the top show that, yes, Mythos isn't too much better than any of the…

Look at those graphs another time. Claude beats gpt.

Can you explain where you're seeing that? From what I see, the first two graphs have OpenAI models above Claude models (including Mythos) on the Technical Non-Expert and the Practitioner evals. Mythos now beats Codex 5.3 on the Expert eval and Opus was already on top for the Apprentice one although now Mythos leads there.

So, even including Mythos, OpenAI still has 2 models on top for the 4 evals listed.

Re: Evaluation of Claude Mythos Preview's cyber capabilities

#26
post #12

So around $10K for a full network takeover with Mythos in 'The Last Ones' (a 32-step simulated corporate network attack). Some limitations from the paper on arxiv (emphasis mine): - No active defenders. Real networks have security teams monitoring for intrusions, responding to alerts, and adapting defences. Our ranges are static, for example our deployment of Elastic Defend was not configured to block or impede attac…

> No active defenders. Real networks have security teams monitoring for intrusions, responding to alerts, and adapting defences.

So, all the fake networks in use outside of Fortune 500?

Re: Evaluation of Claude Mythos Preview's cyber capabilities

#27

Earlier quoted context omitted.

Look at those graphs another time. Claude beats gpt.

Can you explain where you're seeing that? From what I see, the first two graphs have OpenAI models above Claude models (including Mythos) on the Technical Non-Expert and the Practitioner evals. Mythos now beats Codex 5.3 on the Expert eval and Opus was already on top for the Apprentice one although now Mythos leads there. So, even including Mythos, OpenAI still has 2 models on top for the 4 evals listed.

[dead]

Re: Evaluation of Claude Mythos Preview's cyber capabilities

#28
post #12

So around $10K for a full network takeover with Mythos in 'The Last Ones' (a 32-step simulated corporate network attack). Some limitations from the paper on arxiv (emphasis mine): - No active defenders. Real networks have security teams monitoring for intrusions, responding to alerts, and adapting defences. Our ranges are static, for example our deployment of Elastic Defend was not configured to block or impede attac…

> Real networks have security teams monitoring for intrusions, responding to alerts, and adapting defences.

I got some bad news if you really think most (even large) companies have ever once actually looked at what that big Splunk system is collecting for them.

Re: Evaluation of Claude Mythos Preview's cyber capabilities

#29

Earlier quoted context omitted.

Look at those graphs another time. Claude beats gpt.

Can you explain where you're seeing that? From what I see, the first two graphs have OpenAI models above Claude models (including Mythos) on the Technical Non-Expert and the Practitioner evals. Mythos now beats Codex 5.3 on the Expert eval and Opus was already on top for the Apprentice one although now Mythos leads there. So, even including Mythos, OpenAI still has 2 models on top for the 4 evals listed.

> From what I see, the first two graphs have OpenAI models above Claude

That's just in that final graph, and that graph is perhaps the least instructive - they talk about ranges of outcomes but they don't show whether all of the models besides Mythos / Opus 4.6 overlap

Take a look at all three graphs together and it's clear Anthropic are doing better in this arena

Re: Evaluation of Claude Mythos Preview's cyber capabilities

#30
post #28
post #12

So around $10K for a full network takeover with Mythos in 'The Last Ones' (a 32-step simulated corporate network attack). Some limitations from the paper on arxiv (emphasis mine): - No active defenders. Real networks have security teams monitoring for intrusions, responding to alerts, and adapting defences. Our ranges are static, for example our deployment of Elastic Defend was not configured to block or impede attac…

> Real networks have security teams monitoring for intrusions, responding to alerts, and adapting defences. I got some bad news if you really think most (even large) companies have ever once actually looked at what that big Splunk system is collecting for them.

The great irony is that now that Splunk audit trail will probably end up being consumed by LLMs on the lookout for threat actors who are probably also using LLMs to attempt intrusions.

It's a great time to be selling GPUs.

Post reply on HN