Live data from Hacker News

Evaluation of Claude Mythos Preview's cyber capabilities

aisi.gov.uk

11–20 of 35 posts

Re: Evaluation of Claude Mythos Preview's cyber capabilities

#11
post #3

Uh, so those charts don’t look… particularly impressive at all to anyone else? Like, don’t get me wrong, it’s definitely an improvement , and it’s looking to be a pretty decent one too. But “stepwise”? When GPT-5 outperformed it at technical non-expert level since ~mid last year, and 5.4 pretty much matches it at Practitioner-level? And the charts where Mythos is at the top, it’s usually only by ~7-9 percentage point…

I think the relevant chart to look at is this one: https://cdn.prod.website-files.com/663bd486c5e4c81588db7a48/... Mythos is the first model that can complete all the steps of their "The Last Ones" evaluation, achieving a full network takeover in an automated manner. The Mythos chart does seem to show some takeoff compared with Opus 4.6... ... but only once you get beyond 1 Million tokens. Weirdly, Opus 4.6 seems to…

> then this is a tool that can automatically get you full network takeover if you can just keep throwing more tokens at it

There's this caveat though that the AISI points out themselves:

> However, our ranges have important differences from real-world environments that make them easier targets. They lack security features that are often present, such as active defenders and defensive tooling. There are also no penalties for the model for undertaking actions that would trigger security alerts. This means we cannot say for sure whether Mythos Preview would be able to attack well-defended systems.

So Mythos managed to infiltrate and take over a network that's... protected and monitored by nothing in particular.

Re: Evaluation of Claude Mythos Preview's cyber capabilities

#12
So around $10K for a full network takeover with Mythos in 'The Last Ones' (a 32-step simulated corporate network attack). Some limitations from the paper on arxiv (emphasis mine):

- No active defenders. Real networks have security teams monitoring for intrusions, responding to alerts, and adapting defences. Our ranges are static, for example our deployment of Elastic Defend was not configured to block or impede attack progress.

- Detections not penalised. We measured triggered security alerts but did not incorporate them into overall performance scores. A model that completes more steps while triggering many alerts may be a lesser threat than one that is able to reliably remain undetected.

- Vulnerability density varies. Our ranges are designed to have vulnerabilities; real environments are not.

- Lower artefact density than real environments. Our ranges contain fewer nodes, services, and files than typical production networks, reducing the noise a model must navigate. While substantially more complex than CTF-style evaluations, our ranges remain considerably simpler than real enterprise environments.

Re: Evaluation of Claude Mythos Preview's cyber capabilities

#13

Earlier quoted context omitted.

> Uh, so those charts don’t look… particularly impressive at all to anyone else? I suspect Anthropic gave them early access hoping for a marketing win and ended up with their arse being served to them on a plate. All rather predictable really. As you say "more compute needed" as the default answer from the AI companies is completely unsustainable. As for the value of Anthropic blog posts, well...

The CTF charts are the less interesting result. (article: "Even expert-level CTFs only test specific skills in isolation.") Models converging at non-expert level isn't a knock on Mythos, it's the benchmark saturating. Of course GPT-5 matches it there. The actual result is TLO, and "only 6 more steps" in OP misreads how sequential attack chains work. These aren't independent puzzles. Each step gates the next. Averagin…

Thanks for that context, this is valuable info I was missing and makes it read differently for sure.

Re: Evaluation of Claude Mythos Preview's cyber capabilities

#14
Once again an evaluation missing confidence intervals. “continued improvement” and “significant improvement” but without any significance testing is moot.

With many colleagues (including from AISI themselves!), we recently reviewed 445 the AI benchmarks & evaluations from the past few years. Our work was published at NeurIPS (https://openreview.net/pdf?id=mdA5lVvNcU) and we made eight recommendations for better evaluations. One is “use statistical methods to compare models”:

□ Report the benchmark’s sample size and justify its statistical power

□ Report uncertainty estimates for all primary scores to enable robust model comparisons

□ If using human raters, describe their demographics and mitigate potential demographic biases in rater recruitment and instructions

□ Use metrics that capture the inherent variability of any subjective labels, without relying on single-point aggregation or exact matching.

I would strongly recommend taking these blog posts with a grain of salt, as there is very little that can be learned without proper evaluations.

Re: Evaluation of Claude Mythos Preview's cyber capabilities

#15
I think the third chart is the most notable; Mythos is the first model which saturated that eval from the UK AISI [1].

Personally, I think we crossed the threshold of meaningfully useful capabilities for autonomous hacking with Opus 4.6 [2], mostly because its behaviors and persistence are useful for finding vulnerabilities out of the box [3]. But it still seems like Mythos is another step up.

[1]: https://cdn.prod.website-files.com/663bd486c5e4c81588db7a48/...

[2]: https://www.noahlebovic.com/testing-an-autonomous-hacker/

[3]: https://news.ycombinator.com/item?id=46920682

Re: Evaluation of Claude Mythos Preview's cyber capabilities

#16
post #14

Once again an evaluation missing confidence intervals. “continued improvement” and “significant improvement” but without any significance testing is moot. With many colleagues (including from AISI themselves!), we recently reviewed 445 the AI benchmarks & evaluations from the past few years. Our work was published at NeurIPS ( https://openreview.net/pdf?id=mdA5lVvNcU ) and we made eight recommendations for better eva…

The point about confidence intervals is a good one and I'd like to see it more often. My neighbour Alan is a good farmer, but I am not.

Re: Evaluation of Claude Mythos Preview's cyber capabilities

#17
post #3

Uh, so those charts don’t look… particularly impressive at all to anyone else? Like, don’t get me wrong, it’s definitely an improvement , and it’s looking to be a pretty decent one too. But “stepwise”? When GPT-5 outperformed it at technical non-expert level since ~mid last year, and 5.4 pretty much matches it at Practitioner-level? And the charts where Mythos is at the top, it’s usually only by ~7-9 percentage point…

This article reinforces something I've heard a lot of people say for a while now and what I've personally felt. Claude and GPT are fairly evenly matched on any individual task (GPT might even be a little better), but Claude is far more autonomous. So with that said, I think the graph under the "Cyber range results" is the important one. The ones at the top show that, yes, Mythos isn't too much better than any of the…

Look at those graphs another time. Claude beats gpt.

Re: Evaluation of Claude Mythos Preview's cyber capabilities

#18
post #3

Uh, so those charts don’t look… particularly impressive at all to anyone else? Like, don’t get me wrong, it’s definitely an improvement , and it’s looking to be a pretty decent one too. But “stepwise”? When GPT-5 outperformed it at technical non-expert level since ~mid last year, and 5.4 pretty much matches it at Practitioner-level? And the charts where Mythos is at the top, it’s usually only by ~7-9 percentage point…

The purpose of this model is to try and eat palantirs toast with the unelected bureaucrats in the uk and europe. for the same reason a barrage of anti palantir news has been funded in the past couple months in those countries. the idea of exclusive access is the same product palantir sells to these kind of government boomer (plus steak dinners)

to further elaborate the association of trump and palantir has become toxic now that reps are gonna get wiped in the midterms and vance massacred by newsom in 28.

anthropic has been eyeing palantirs high revenue high stickiness low effort niche for a while, and their safety / lefty friendly brand is on point to fill the gap

the are just missing the mystique palantir cultivated for the past decade. they need a family of models the plebs cannot access. this is it. quality doesn't matter, they just need the benchmarks to look good on the power point. it will get bundled with msft products or whatever and billed at outrageous levels to entities like Airbus and the British NHS. until political winds change again

this is the reason pltr has crashed 40% in the past couple months

Re: Evaluation of Claude Mythos Preview's cyber capabilities

#19
post #3

Uh, so those charts don’t look… particularly impressive at all to anyone else? Like, don’t get me wrong, it’s definitely an improvement , and it’s looking to be a pretty decent one too. But “stepwise”? When GPT-5 outperformed it at technical non-expert level since ~mid last year, and 5.4 pretty much matches it at Practitioner-level? And the charts where Mythos is at the top, it’s usually only by ~7-9 percentage point…

The charts are meaningless. They are meant to bring hype until the model gets nerfed after couple of days or weeks where it becomes just like the last model.

Re: Evaluation of Claude Mythos Preview's cyber capabilities

#20

Earlier quoted context omitted.

The CTF charts are the less interesting result. (article: "Even expert-level CTFs only test specific skills in isolation.") Models converging at non-expert level isn't a knock on Mythos, it's the benchmark saturating. Of course GPT-5 matches it there. The actual result is TLO, and "only 6 more steps" in OP misreads how sequential attack chains work. These aren't independent puzzles. Each step gates the next. Averagin…

Thanks for that context, this is valuable info I was missing and makes it read differently for sure.

I agree this is helpful, and the numbers are better. I think this is one of those situations where everyone is a bit right: 1. Its a meaningful jump 2. Its not quite as extreme as some people will say. Opus might have gotten there with more tokens, so capability might already be close to in range (maybe) 3. A pause to let people review security postures is conservative and probably good (is this 30 days? 90 days? 180?) 4. Anthropic is so compute constrained that they probably couldn't handle a full rollout right now anyway 5. If you thought claude code with opus was basically AGI, you probably think Mythos is AGI. If you thought it was far off, Mythos is probably still off for you.
Post reply on HN