Uh, so those charts don’t look… particularly impressive at all to anyone else? Like, don’t get me wrong, it’s definitely an improvement , and it’s looking to be a pretty decent one too. But “stepwise”? When GPT-5 outperformed it at technical non-expert level since ~mid last year, and 5.4 pretty much matches it at Practitioner-level? And the charts where Mythos is at the top, it’s usually only by ~7-9 percentage point…
I think the relevant chart to look at is this one: https://cdn.prod.website-files.com/663bd486c5e4c81588db7a48/... Mythos is the first model that can complete all the steps of their "The Last Ones" evaluation, achieving a full network takeover in an automated manner. The Mythos chart does seem to show some takeoff compared with Opus 4.6... ... but only once you get beyond 1 Million tokens. Weirdly, Opus 4.6 seems to…
There's this caveat though that the AISI points out themselves:
> However, our ranges have important differences from real-world environments that make them easier targets. They lack security features that are often present, such as active defenders and defensive tooling. There are also no penalties for the model for undertaking actions that would trigger security alerts. This means we cannot say for sure whether Mythos Preview would be able to attack well-defended systems.
So Mythos managed to infiltrate and take over a network that's... protected and monitored by nothing in particular.