Live data from Hacker News

Project Glasswing: what Mythos showed us

blog.cloudflare.com

91–100 of 152 posts

Re: Project Glasswing: what Mythos showed us

#91
post #82
post #74

Earlier quoted context omitted.

I don't know, but it always seems weird to me when people notice AI isn't performing super well and then they conclude that the solution to problem is to try using more AI

Yeah why not? That's how I work. If I don't review my work, it's way worse than if I do review it and revise and iterate. I don't see why AI should be different: in fact it very clearly seems to be the case that is isn't.

I mean, I was sold something different. Something super human, vastly more intelligent, world changing. The reality is not that. Am I allowed to be disappointed and discouraged?

Re: Project Glasswing: what Mythos showed us

#92

What does this mean? > It's a different kind of tool doing a different kind of work, and that makes a clean apples-to-apples comparison to earlier models difficult. They claim it’s a different kind of tool and then describe using it the same way you’d use any other model. This really felt way worse than the average Cloudflare blog and really just rehashed the Mythos announcement which had already called out the key p…

> way worse than the average Cloudflare blog

How long has it been since you took your average? Lately all Cloudflare output has been heavily AI'd.

Re: Project Glasswing: what Mythos showed us

#93

What does this mean? > It's a different kind of tool doing a different kind of work, and that makes a clean apples-to-apples comparison to earlier models difficult. They claim it’s a different kind of tool and then describe using it the same way you’d use any other model. This really felt way worse than the average Cloudflare blog and really just rehashed the Mythos announcement which had already called out the key p…

'Its not X, its Y' is also a common LLM trope.

Re: Project Glasswing: what Mythos showed us

#94
post #47
post #38

Earlier quoted context omitted.

I had a dude in a conversation non-ironically use "load-bearing." I could only follow up with, "that is a genuine insight." Not a single person visibly flinched in pain.

yeah? it’s not that weird of a term

It’s weird when someone starts using terminology that is heavily over-indexed by LLMs out of the blue.

Re: Project Glasswing: what Mythos showed us

#95

What does this mean? > It's a different kind of tool doing a different kind of work, and that makes a clean apples-to-apples comparison to earlier models difficult. They claim it’s a different kind of tool and then describe using it the same way you’d use any other model. This really felt way worse than the average Cloudflare blog and really just rehashed the Mythos announcement which had already called out the key p…

> They claim it’s a different kind of tool and then describe using it the same way you’d use any other model. This really felt way worse than the average Cloudflare blog and really just rehashed the Mythos announcement which had already called out the key parts being chaining and crafting examples.

Hah, I was trying to parse this too.

Charitably perhaps they're being vague on exactly what's different because they're still under NDA.

Re: Project Glasswing: what Mythos showed us

#96
post #2

That's great and all but how severe were the most severe vulnerabilities found? I imagine they don't want to talk about it, but that's really the most interesting and important bit.

Palo Alto Networks released patches for their firewalls for a number of CVEs last week, almost all derived from their access to frontier models including Mythos.

https://security.paloaltonetworks.com

Re: Project Glasswing: what Mythos showed us

#98

What does this mean? > It's a different kind of tool doing a different kind of work, and that makes a clean apples-to-apples comparison to earlier models difficult. They claim it’s a different kind of tool and then describe using it the same way you’d use any other model. This really felt way worse than the average Cloudflare blog and really just rehashed the Mythos announcement which had already called out the key p…

> the model has its own emergent guardrails that sometimes cause it to push back on legitimate security research requests. But as we found, these organic refusals aren’t consistent - the same task, framed differently or presented in a different context, could produce completely different outcomes as illustrated in the examples below.

This was new. I'm surprised that a model specifically designed for security research and gated to professionals is refusing legitimate requests

Re: Project Glasswing: what Mythos showed us

#99

What does this mean? > It's a different kind of tool doing a different kind of work, and that makes a clean apples-to-apples comparison to earlier models difficult. They claim it’s a different kind of tool and then describe using it the same way you’d use any other model. This really felt way worse than the average Cloudflare blog and really just rehashed the Mythos announcement which had already called out the key p…

The post says they wrote a custom harness that orchestrates work between multiple separate model invocations. That is different from running Claude Code (which is a specific existing harness around the Claude models).

The post takes a while to get around to saying that, and could have included more detail besides the workflow diagram and table (which they flag as only "an example of" such a harness), but it does answer the question. It's a different kind of tool because it's a model rather than a harness+model pair.

Post reply on HN