Live data from Hacker News

Small models also found the vulnerabilities that Mythos found

aisle.com

191–200 of 372 posts

Re: Small models also found the vulnerabilities that Mythos found

#191

Earlier quoted context omitted.

No, they didn't. They distinguished it, when presented with it. Wildly different problem.

Yeah. And it is totally depressing that this article got voted to the top of the front page. It means people aren’t capable of this most basic reasoning so they jumped on the “aha! so the mythos announcement was just marketing!!”

[dead]

Re: Small models also found the vulnerabilities that Mythos found

#192
post #124

The best way to think of Anthropic's communication about Mythos is as advertisement. It's basically "our model is too smart to release" which suggests they're ahead of OpenAI (without proof)

Seen similar things with Openai and Palantir.

Yes. OpenAI does the exact same thing.

Re: Small models also found the vulnerabilities that Mythos found

#193

Earlier quoted context omitted.

No, they didn't. They distinguished it, when presented with it. Wildly different problem.

Yeah. And it is totally depressing that this article got voted to the top of the front page. It means people aren’t capable of this most basic reasoning so they jumped on the “aha! so the mythos announcement was just marketing!!”

Yeah. Extremely disappointing.

Re: Small models also found the vulnerabilities that Mythos found

#194
post #171

Earlier quoted context omitted.

> because small models found the same vulnerability. With a ton of extra support. Note this key passage: >We isolated the vulnerable svc_rpc_gss_validate function, provided architectural context (that it handles network-parsed RPC credentials, that oa_length comes from the packet), and asked eight models to assess it for security vulnerabilities. Yeah it can find a needle in a haystack without false positives, if you…

The benefit here is reducing the time to find vulnerabilities; faster than humans, right? So if you can rig a harness for each function in the system, by first finding where it’s used, its expected input, etc, and doing that for all functions, does it discover vulnerabilities faster than humans? Doesn’t matter that they isolated one thing. It matters that the context they provided was discoverable by the model.

There is absolutely zero reason to believe you could use this same approach to find and exploit vulns without Mythos finding them first. We already know that older LLMs can’t do what Mythos has done. Anthropic and others have been trying for years.

Re: Small models also found the vulnerabilities that Mythos found

#195
I feel like there have been enough hyperbolic claims by Anthropic, that I'm starting to get some real Boy Who Cried Wolf energy. I'm starting to tune out, and assume it is a marketing ploy. Trust me, I'm an Antropic fan, and I pay my $200/month for max, but the claims are wearing thin.

Re: Small models also found the vulnerabilities that Mythos found

#197
post #164
post #155

Earlier quoted context omitted.

Also, what is $20,000 today can be $2000 next year. Or $20... See e.g. https://epoch.ai/data-insights/llm-inference-price-trends/

Or $200,000 for consumers when they have to make a profit

Good point. This is why consumer phones have got much worse since 2005 and now cost millions of dollars.

Re: Small models also found the vulnerabilities that Mythos found

#198

So there are two competing narratives: 1. Mythos uniquely is able to find vulnerabilities that other LLMs cannot practically. 2. All LLMs could already do this but no one tried the way anthropic did. The truth is one of these. And it comes down whether the comparison is apples to apples. Since we don't know the exact specifics of how either tests were performed, we lack a way of knowing absolutely. So I guess, like s…

People have found 0days assisted by LLMs for a while, and none of them wrote hype pieces to find an excuse not to release their 10x bigger model in the middle of a GPU shortage.

https://sean.heelan.io/2025/05/22/how-i-used-o3-to-find-cve-...

Re: Small models also found the vulnerabilities that Mythos found

#200
post #130

Earlier quoted context omitted.

The citation is the Anthropic writeup.

They did not say what you are saying… > If you try to automate a small model to look for vulnerabilities over 10,000 files, it's going to say there are 9,500 vulns.

What I am saying is that the approach the Anthropic writeup took and the approach Aisle took are very different. The Aisle approach is vastly easier on the LLM. I don't think I need a citation for that. You can just read both writeups.

The "9500" quote is my conjecture of what might happen if they fix their approach, but the burden of proof is definitely not on me to actually fix their writeup and spend a bunch of money to run a new eval! They are the ones making a claim on shaky ground, not me.

Post reply on HN