Live data from Hacker News

Project Glasswing: what Mythos showed us

blog.cloudflare.com

61–70 of 152 posts

Re: Project Glasswing: what Mythos showed us

#62

What does this mean? > It's a different kind of tool doing a different kind of work, and that makes a clean apples-to-apples comparison to earlier models difficult. They claim it’s a different kind of tool and then describe using it the same way you’d use any other model. This really felt way worse than the average Cloudflare blog and really just rehashed the Mythos announcement which had already called out the key p…

I think what they might mean is: Because of it's capabilities, a new kind of harness can be built for it, thus the entire system (model + harness) is a different kind of tool than say Claude code

But did they build this different harness? And are they sure other models can't cope with it?

Re: Project Glasswing: what Mythos showed us

#63
post #53

Earlier quoted context omitted.

But why would cloudflare advertise Anthropic? They are competing with Anthropic by hosting open weights models.

https://www.cloudflare.com/press/press-releases/2025/cloudfl...

It's circular financing. It's a circlejerk. It's a circular financing jerk.

Re: Project Glasswing: what Mythos showed us

#64
post #40

Earlier quoted context omitted.

In the last few days I was recommending to read the insights from XBOW [1], it's a competitor but it adds more information to the discussion. [1] https://xbow.com/blog/mythos-offensive-security-xbow-evaluat...

Thanks for sharing. Its definitely more concrete. Some of the things that I was hoping to find were, the number of false positives, the times it takes to identify the false positives from real ones, the taxation on human mind to perform this exercise. Did anyone manually verified the exploits which were identified by the LLM or were they assumed correct based on the explanation. I do understand that the target audien…

>, the number of false positives,

Really this is why the LLM needs to be able to write exploits for issues it finds. Of course that leads down a rabbit hole of other issues. But if an exploit works, then that's pretty conclusive evidence.

Re: Project Glasswing: what Mythos showed us

#65
post #32

Earlier quoted context omitted.

That's a scary thought, llm's training on llm output. People trained by default of ubiquity to think and read llm output produce their own llm-esque writing. Seems stifling. We'll need someway to reward human creativity and out-of-bounds thinking before our greatest corpus of human intellect is a bounded by whenever and whatever was trained on.

Writing and later the printing press have already considerably stifled human expressiveness. Language used to be noch more fragmented and diverse before mass media (or the Bible in every household). In my grandmother’s time you would have difficulty understanding people from three villages down the road.

I'm not sure enabling people three villages apart to communicate with each other counts as "stifling human expressiveness"

Re: Project Glasswing: what Mythos showed us

#66
post #62

Earlier quoted context omitted.

I think what they might mean is: Because of it's capabilities, a new kind of harness can be built for it, thus the entire system (model + harness) is a different kind of tool than say Claude code

But did they build this different harness? And are they sure other models can't cope with it?

Right I expected the piece to transition into “and here’s how we built a whole new thing for it” but it never did.

Re: Project Glasswing: what Mythos showed us

#67
post #3

The real question is whether it was Mythos or Opus that wrote this post. > "Why it matters" It doesn't, it's a corporate blog, they were rarely written in one-author's voice anyway, but it's interesting to see that even large organisations are outsourcing their blogs to LLMs.

Disappointing really.

Re: Project Glasswing: what Mythos showed us

#69
post #64

Earlier quoted context omitted.

Thanks for sharing. Its definitely more concrete. Some of the things that I was hoping to find were, the number of false positives, the times it takes to identify the false positives from real ones, the taxation on human mind to perform this exercise. Did anyone manually verified the exploits which were identified by the LLM or were they assumed correct based on the explanation. I do understand that the target audien…

>, the number of false positives, Really this is why the LLM needs to be able to write exploits for issues it finds. Of course that leads down a rabbit hole of other issues. But if an exploit works, then that's pretty conclusive evidence.

For a subset of bugs, yes. For some others, not really: I've seen LLMs make bogus assumptions about the threat model (in which case, the exploit works but doesn't demonstrate anything useful) or "cheat" by modifying the code to demonstrate a hallucinated issue.

Frontier models, including Mythos, can greatly streamline bug hunting and exploit developments in the hands of a competent security engineer. In the hands of a person with no security experience, they will still mostly waste your time and money.

Re: Project Glasswing: what Mythos showed us

#70
post #12
post #2

That's great and all but how severe were the most severe vulnerabilities found? I imagine they don't want to talk about it, but that's really the most interesting and important bit.

As much as I’d like to share in the skepticism, the very beginning of the article states it very plainly — this is a step function. Lots of people feel that Mythos is a psyops campaign, but I don’t really understand the skepticism. Most of it seems to stem from the general distrust of things that aren’t publicly available. A few Anthropic employees have described Mythos as a general purpose model improvement, but tha…

> As much as I’d like to share in the skepticism, the very beginning of the article states it very plainly — this is a step function.

To be fair, they can't say "You know, Mythos is better, but improvements are overhyped af". Moreover, their explanation of that "step change" is strange. It sounds like Mythos isn't that much better at finding vulnerabilities (which is very strange, given statements from Mozilla), but is way stronger at working with them.

> Lots of people feel that Mythos is a psyops campaign, but I don’t really understand the skepticism. Most of it seems to stem from the general distrust of things that aren’t publicly available.

1) Attempts to spin the idea about "Super powerful general purpose model that can't be released for some not so clear reasons" are usually a very bad sign. OpenAI proves it.

2) Mythos system card has a lot of strange moments, errors and things that sound like attempts to deceive.

3) It's strange that Anthropic is struggling with both Sonnet 5.0 and Opus 5.0, but at the same time has a breakthrough in the form of Mythos.

> A few Anthropic employees have described Mythos as a general purpose model improvement, but that claim has yet to be widely backed up so that’s the only place I’m remaining skeptical.

Article describes Mythos as a cybersecurity-specific model though. It's yet another unclear moment.

Post reply on HN