Live data from Hacker News

Project Glasswing: Securing critical software for the AI era

anthropic.com

451–460 of 921 posts

Re: Project Glasswing: Securing critical software for the AI era

#451

> Mythos Preview identified a number of Linux kernel vulnerabilities that allow an adversary to write out-of-bounds (e.g., through a buffer overflow, use-after-free, or double-free vulnerability.) Many of these were remotely-triggerable. However, even after several thousand scans over the repository, because of the Linux kernel’s defense in depth measures Mythos Preview was unable to successfully exploit any of these…

Kernel address space layout randomization they are talking about is a bit different than (x != null). Other bug may allow to locate the required address.

Re: Project Glasswing: Securing critical software for the AI era

#452

> On the global stage, state-sponsored attacks from actors like China, Iran, North Korea, and Russia have threatened to compromise the infrastructure that underpins both civilian life and military readiness. AITA for thinking that PRISM was probably the state sponsored program affecting civilian life the most? And that one state is missing from the list here?

Look, we have always been at war with EastAsia.

Re: Project Glasswing: Securing critical software for the AI era

#453

Earlier quoted context omitted.

Or more likely, its just an exaggeration or lie.

Yes I'm sure this is all a massive conspiracy by the many companies that are making statements alongside Anthropic

[dead]

Re: Project Glasswing: Securing critical software for the AI era

#454

Now, its very possible that this is Anthropic marketing puffery, but even if it is half true it still represents an incredible advancement in hunting vulnerabilities. It will be interesting to see where this goes. If its actually this good, and Apple and Google apply it to their mobile OS codebases, it could wipe out the commercial spyware industry, forcing them to rely more on hacking humans rather than hacking mobi…

Its not, if you dont trust Anthropic, I hope you trust Daniel Steinberg of curl, who has said AI has gotten really good at detecting bugs and vulnerabilities. Here is his LinkedIN post https://www.linkedin.com/posts/danielstenberg_hackerone-acti...

Didn’t they ban issues generated by ai?

Re: Project Glasswing: Securing critical software for the AI era

#455
post #206

Earlier quoted context omitted.

I don't trust a corpo to choose what is "most critical". That's what's messed up about it.

That is a fine stance to hold but some facts are still true regardless of your view on large businesses. For example, it will benefit more people to secure Microsoft or Amazon services than it would be to secure a smaller, less corporate player in those same service ecosystems. You could go on to argue that the second order effects of improving one service provider over another chooses who gets to play, but that is t…

The longer term economic outcome of this is consolidation: large players get stronger, weak players get weaker.

That's not a good outcome for the economy.

Re: Project Glasswing: Securing critical software for the AI era

#456

I think this is a largely inflated PR stunt. Opus 4.6 was already capable of finding 0days and chaining together vulns to create exploits. See [0] and [1]. [0] https://www.csoonline.com/article/4153288/vim-and-gnu-emacs-... [1] https://xbow.com/blog/top-1-how-xbow-did-it

Did you read the article?

Re: Project Glasswing: Securing critical software for the AI era

#457
post #17

One of the things I'm always looking at with new models released is long context performance, and based on the system card it seems like they've cracked it: GraphWalks BFS 256K-1M Mythos Opus GPT5.4 80.0% 38.7% 21.4%

Huh, I don’t know what “long context performance” means exactly in these tests, so completely anecdotally , my experience with gpt5.4 via codex cli vs Claude code opus, gpt5.4 seems to do significantly better in long contexts I think partly due to some special context compaction stored in encrypted blobs. On long conversations opus in Claude code will for me lose memory of what we were working on earlier, whereas one…

This isn’t talking about compaction. This refers to performance as the model is loaded with 500k to 1m tokens.

Re: Project Glasswing: Securing critical software for the AI era

#458

To be clear, we don’t know that this tool is better at finding bugs than fuzzing. We just know that it’s finding bugs that fuzzing missed. It’s possible fuzzing also finds bugs that this AI would miss.

This is obviously just cope (there's a long, strong-form argument for why LLM-agent vulnerability research is plausibly much more potent than fuzzing, but we don't have to reach it because you can dispose of the whole argument by noting that agents can build and drive fuzzers and triage their outputs), but what I'd really like to understand better is why? What's the impetus to come up with these weird rationalizations for why it's not a big deal that frontier models can identify bugs everyone else missed and then construct exploits for them?

Re: Project Glasswing: Securing critical software for the AI era

#459

> Mythos Preview identified a number of Linux kernel vulnerabilities that allow an adversary to write out-of-bounds (e.g., through a buffer overflow, use-after-free, or double-free vulnerability.) Many of these were remotely-triggerable. However, even after several thousand scans over the repository, because of the Linux kernel’s defense in depth measures Mythos Preview was unable to successfully exploit any of these…

> The model autonomously found and chained together several vulnerabilities in the Linux kernel—the software that runs most of the world’s servers—to allow an attacker to escalate from ordinary user access to complete control of the machine.

Re: Project Glasswing: Securing critical software for the AI era

#460
post #381

Earlier quoted context omitted.

are we cooked yet? Benchmarks look very impressive! even if they're flawed, it still translates to real world improvements

People say we're cooked every single day. The only response is to continue life as if we aren't. When we are, you won't have to ask that question.

Everyone’s pretending the suits are going to want to do the prompting. We all know they aren’t.
Post reply on HN