Live data from Hacker News

Assessing Claude Mythos Preview's cybersecurity capabilities

red.anthropic.com

21–30 of 59 posts

Re: Assessing Claude Mythos Preview's cybersecurity capabilities

#21
The elephant in the room here is that there are hundreds of millions of embedded devices that cannot be upgraded easily and will be running vulnerable binaries essentially forever. This was a problem before of course, but the ease of chaining vulnerabilities takes the issue to a new level.

The only practical defense is for these frontier models to generate _beneficial_ attacks to innoculate older binaries by remote exploits. I dubbed these 'antibotty' networks in a speculative paper last year, but never thought things would move this fast! https://anil.recoil.org/papers/2025-internet-ecology.pdf

Re: Assessing Claude Mythos Preview's cybersecurity capabilities

#23
post #21

The elephant in the room here is that there are hundreds of millions of embedded devices that cannot be upgraded easily and will be running vulnerable binaries essentially forever. This was a problem before of course, but the ease of chaining vulnerabilities takes the issue to a new level. The only practical defense is for these frontier models to generate _beneficial_ attacks to innoculate older binaries by remote e…

No, the elephant in the room is that even bad actors will now have easier to find vulnerabilities in, maintained or not, widely or in critical places used software. Unmaintained and remotely accessible devices should be discarded as soon as possible, you can't stay waiting till some of the good guys decide to give some time to your niche but critical unmaintained piece of software. Because if there is a possibility of taking profit of it, it will be checked and exploited.

And you can't assume that whatever vulnerability they have will let good guys to do the extra (and legally risky) work of closing the hole.

Re: Assessing Claude Mythos Preview's cybersecurity capabilities

#24
post #20

Related ongoing threads: System Card: Claude Mythos Preview [pdf] - https://news.ycombinator.com/item?id=47679258 Project Glasswing: Securing critical software for the AI era - https://news.ycombinator.com/item?id=47679121 I can't tell which of the current threads, if any, should be merged - they all seem significant. Anyone?

There is a lot to digest here. Maybe having a few separate pages makes them a bit more digestible. The system card itself is some 200 odd pages

Re: Assessing Claude Mythos Preview's cybersecurity capabilities

#25

This is becoming a bit scary. I almost hope we'll reach some kind of plateau for llm intelligence soon.

The immediate plateau is the energy output of the Sun captured by the Dyson Swarm around it. Until there it's smooth sailing.

Re: Assessing Claude Mythos Preview's cybersecurity capabilities

#26
post #7

Earlier quoted context omitted.

If we don't innovate, someone else will. This is the very nature of being a human being. We summit mountains, regardless of the danger or challenge.

>If we don't innovate, someone else will. Terrible take. You don't get to push the extinction button just because you think China will beat you to the punch. >This is the very nature of being a human being. We summit mountains, regardless of the danger or challenge. No, just no... We barely survived the Cold War, at times because of pure luck. AI is at least as dangerous as that, if not more. We have far exceeded our…

You assume there is the option of not pushing the extinction button. Nobody asked chimps if they wanted humans around. This processes are outside control.

Re: Assessing Claude Mythos Preview's cybersecurity capabilities

#27

Since this level of security ”scanning” requires heaps of money, this is going to kill off a substantial part of F/OSS.

Well, maybe not... see Simon Willison's ongoing reporting [0] on all the bug reports for `curl` people are finding with LLMs.

Interesting to see them go from "DON'T GIVE US AI SLOP!" to "Wow, lots of actual bugs found, including [ed: at least one] bug found by two people!"

[0]: https://simonwillison.net/search/?q=curl

Re: Assessing Claude Mythos Preview's cybersecurity capabilities

#28
post #21

The elephant in the room here is that there are hundreds of millions of embedded devices that cannot be upgraded easily and will be running vulnerable binaries essentially forever. This was a problem before of course, but the ease of chaining vulnerabilities takes the issue to a new level. The only practical defense is for these frontier models to generate _beneficial_ attacks to innoculate older binaries by remote e…

No, the elephant in the room is that even bad actors will now have easier to find vulnerabilities in, maintained or not, widely or in critical places used software. Unmaintained and remotely accessible devices should be discarded as soon as possible, you can't stay waiting till some of the good guys decide to give some time to your niche but critical unmaintained piece of software. Because if there is a possibility o…

_SHOULD_ yes sure, but realistically is that going to happen?

Re: Assessing Claude Mythos Preview's cybersecurity capabilities

#29

Earlier quoted context omitted.

No, the elephant in the room is that even bad actors will now have easier to find vulnerabilities in, maintained or not, widely or in critical places used software. Unmaintained and remotely accessible devices should be discarded as soon as possible, you can't stay waiting till some of the good guys decide to give some time to your niche but critical unmaintained piece of software. Because if there is a possibility o…

_SHOULD_ yes sure, but realistically is that going to happen?

As doom and gloom as things are generally, I do think things have gotten better. Due to legislation and commercial pressure things like wifi routers shipping with the same default password and open settings have gotten better. Webhosts and ISPs have implemented many improvements to protecting their residential customers.

I take your point, but think that it's also maybe too far.

Re: Assessing Claude Mythos Preview's cybersecurity capabilities

#30
My two cents is LLMs are way stronger in areas where the reward function is well known, such as exploiting - you break the security, you succeed.

It's much harder to establish whats a usable and well architected, novel piece of software, thus in that area, progress isn't nearly as fast, while here you can just gradient descent your way to world domination, provided you have enough GPUs.

Post reply on HN