Live data from Hacker News

Project Glasswing: An Initial Update

anthropic.com

11–20 of 345 posts

Re: Project Glasswing: An Initial Update

#11

The vulnerabilities found continues to impress, and make legacy media, Twitter and Youtube go nuts. But we still have no data to prove this wasn't doable with the same initiative backed by Opus 4.7, and there is no GA for Mythos access.

> Mozilla found and fixed 271 vulnerabilities in Firefox 150 while testing Mythos Preview—over ten times more than they found in Firefox 148 with Claude Opus 4.6 4.6 but close.

Right, but were they using the same methodology and harness? I'm skeptical that they're doing something with the harness - i.e. with Mythos, they pass each file in one at a time, whereas on 4.6 they let Claude Code run loose to find bugs. This would have a larger impact difference than the model itself.

Re: Project Glasswing: An Initial Update

#12

The vulnerabilities found continues to impress, and make legacy media, Twitter and Youtube go nuts. But we still have no data to prove this wasn't doable with the same initiative backed by Opus 4.7, and there is no GA for Mythos access.

I've seen a blog post by a security researcher saying that he was able to find the same vulnerabilities (for Firefox IIRC) with a ~30B params LLM... So yeah, huge marketing as always.

To me it’s clear what’s going on.

The American firms are focused on marketing now to convince people to not even consider open sourced models / open weight models as they are inferior (that’s what they want you to believe).

Re: Project Glasswing: An Initial Update

#13

The vulnerabilities found continues to impress, and make legacy media, Twitter and Youtube go nuts. But we still have no data to prove this wasn't doable with the same initiative backed by Opus 4.7, and there is no GA for Mythos access.

you would likely be quite interested in the more quantitative writeup from a real research team ! it’s linked about midway in to the article - similar functionally can be reached, yes, but not always and never with fewer tokens than what mythos requires. https://xbow.com/blog/mythos-offensive-security-xbow-evaluat...

Ok this is actually a pretty good article and justifies the step function marketing in security they talked about

Re: Project Glasswing: An Initial Update

#14
post #12

Earlier quoted context omitted.

I've seen a blog post by a security researcher saying that he was able to find the same vulnerabilities (for Firefox IIRC) with a ~30B params LLM... So yeah, huge marketing as always.

To me it’s clear what’s going on. The American firms are focused on marketing now to convince people to not even consider open sourced models / open weight models as they are inferior (that’s what they want you to believe).

IPO is coming is what is going on

Re: Project Glasswing: An Initial Update

#15

The vulnerabilities found continues to impress, and make legacy media, Twitter and Youtube go nuts. But we still have no data to prove this wasn't doable with the same initiative backed by Opus 4.7, and there is no GA for Mythos access.

Makes me wonder if Anthropic is really having issues with allocating compute (see recent deals with xAI and SpaceX). From available benchmarks, it seems like similar results should be possible with GPT 5.5 Pro or Opus 4.7 (with specific cybersecurity trained models).

Who knows but from a valuation stand point it’s better to signal that demand is higher than existing capacity..

Re: Project Glasswing: An Initial Update

#16
post #12

Earlier quoted context omitted.

To me it’s clear what’s going on. The American firms are focused on marketing now to convince people to not even consider open sourced models / open weight models as they are inferior (that’s what they want you to believe).

IPO is coming is what is going on

That’s implicit in my post.

If people actually believe the narrative then the bankers will over price Anthropic and get away with it.

Re: Project Glasswing: An Initial Update

#17

[edit: TFA addresses this, though I still find crazy 90% accuracy overall vs 20% accuracy for curl] Is this suspected vulns or actual vulns? If I recall correctly, it produced 5 for curl but only 1 was legit

I don't know why you're getting downvoted. This is exactly what was reported by curl's creator under the section "Five findings became one": https://daniel.haxx.se/blog/2026/05/11/mythos-finds-a-curl-v...

Re: Project Glasswing: An Initial Update

#18

[edit: TFA addresses this, though I still find crazy 90% accuracy overall vs 20% accuracy for curl] Is this suspected vulns or actual vulns? If I recall correctly, it produced 5 for curl but only 1 was legit

I don't know why you're getting downvoted. This is exactly what was reported by curl's creator under the section "Five findings became one": https://daniel.haxx.se/blog/2026/05/11/mythos-finds-a-curl-v...

[flagged]

Re: Project Glasswing: An Initial Update

#19

The vulnerabilities found continues to impress, and make legacy media, Twitter and Youtube go nuts. But we still have no data to prove this wasn't doable with the same initiative backed by Opus 4.7, and there is no GA for Mythos access.

Makes me wonder if Anthropic is really having issues with allocating compute (see recent deals with xAI and SpaceX). From available benchmarks, it seems like similar results should be possible with GPT 5.5 Pro or Opus 4.7 (with specific cybersecurity trained models).

At least according to this, GPT-5.5 Cyber is on par with Mythic, as the only two models that were able to finish their 32-step corporate network attack simulation.

https://www.aisi.gov.uk/blog/our-evaluation-of-openais-gpt-5...

Re: Project Glasswing: An Initial Update

#20

The vulnerabilities found continues to impress, and make legacy media, Twitter and Youtube go nuts. But we still have no data to prove this wasn't doable with the same initiative backed by Opus 4.7, and there is no GA for Mythos access.

I've seen a blog post by a security researcher saying that he was able to find the same vulnerabilities (for Firefox IIRC) with a ~30B params LLM... So yeah, huge marketing as always.

Did the security researcher point the LLM at the blob of information and say "Find vulnerabilities" or was the LLM told to "determine if vulnerability X is present in this blob"? Confirmation of suspected vulnerabilities is a different problem from finding vulnerabilities.
Post reply on HN