Live data from Hacker News

Project Glasswing: what Mythos showed us

blog.cloudflare.com

81–90 of 152 posts

Re: Project Glasswing: what Mythos showed us

#81
post #8

Beside the poorly written post, the vulnerability discovery workflow might actually give good results

The part on the harness is spot on.

I have been encouraging people to think about agentic coding in the same way.

Let agents do the reading and writing and inspections. Human does the thinking.

Asking an agent that is looking at a firearm specification schematic "what is wrong with this?" and the response is "this thing contains an explosion and can kill". Human "that's the function" when the human should be asking "based upon the materials used, are the fault tolerances sufficient to maintain structural integrity".

Re: Project Glasswing: what Mythos showed us

#82
post #74

> The loudest reaction to Mythos Preview from other security leaders has been about speed - scan faster, patch faster, compress the response cycle. More than one team we have spoken with is now operating under a two-hour SLA from CVE release to patch in production [...] If regression testing takes a day, you cannot get to a two-hour SLA without skipping it, and the bugs you ship when you skip regression testing tend…

I don't know, but it always seems weird to me when people notice AI isn't performing super well and then they conclude that the solution to problem is to try using more AI

Yeah why not? That's how I work. If I don't review my work, it's way worse than if I do review it and revise and iterate. I don't see why AI should be different: in fact it very clearly seems to be the case that is isn't.

Re: Project Glasswing: what Mythos showed us

#83

What does this mean? > It's a different kind of tool doing a different kind of work, and that makes a clean apples-to-apples comparison to earlier models difficult. They claim it’s a different kind of tool and then describe using it the same way you’d use any other model. This really felt way worse than the average Cloudflare blog and really just rehashed the Mythos announcement which had already called out the key p…

My guess is because it is a model trained specifically for security/hacking. So comparing it to Opus, trained for chat/code/etc., is apples-to-oranges.

Re: Project Glasswing: what Mythos showed us

#84
post #38

Earlier quoted context omitted.

I had a dude in a conversation non-ironically use "load-bearing." I could only follow up with, "that is a genuine insight." Not a single person visibly flinched in pain.

Careful, you might have been talking to a Real Engineer. Perhaps even a structural variant that use this phrase pretty much daily.

We weren't talking about "seeing a man about a horse barn" we were talking about software.

Re: Project Glasswing: what Mythos showed us

#85
The "Four lessons" that came out of running this work at scale made me chuckle. Three of the four were essentially identical and entirely obvious. In short: specific, narrow requests work better than "find vulnerabilities." Well, d'uh.

But, I did think the adversarial review (while not novel at all and talked about much in HN circles) is interesting and distinct, at least. I need to put this to work in more of workflows. I think it could be beneficial for non-coding tasks, too.

https://blog.cloudflare.com/cyber-frontier-models/#what-a-ha...

Re: Project Glasswing: what Mythos showed us

#86
post #12
post #2

That's great and all but how severe were the most severe vulnerabilities found? I imagine they don't want to talk about it, but that's really the most interesting and important bit.

As much as I’d like to share in the skepticism, the very beginning of the article states it very plainly — this is a step function. Lots of people feel that Mythos is a psyops campaign, but I don’t really understand the skepticism. Most of it seems to stem from the general distrust of things that aren’t publicly available. A few Anthropic employees have described Mythos as a general purpose model improvement, but tha…

> As much as I’d like to share in the skepticism, the very beginning of the article states it very plainly — this is a step function.

That's great and all, but nobody was being skeptical or asking anything about whether Mythos is or isn't a step function. Mythos could be a ten-dimensional ladder and it wouldn't change my question. The question wasn't about Mythos, but about Cloudflare: what did they found? That question is entirely fair and expected regardless of whether vulnerabilities are found via Mythos, the NSA, or a caveman.

Re: Project Glasswing: what Mythos showed us

#88
post #62

Earlier quoted context omitted.

But did they build this different harness? And are they sure other models can't cope with it?

Right I expected the piece to transition into “and here’s how we built a whole new thing for it” but it never did.

They kind of have a little diagram explaining the steps I imagine every single step in that to basically be it's own Claude code session.

Re: Project Glasswing: what Mythos showed us

#89
post #80

Earlier quoted context omitted.

That's a scary thought, llm's training on llm output. People trained by default of ubiquity to think and read llm output produce their own llm-esque writing. Seems stifling. We'll need someway to reward human creativity and out-of-bounds thinking before our greatest corpus of human intellect is a bounded by whenever and whatever was trained on.

I don't understand this mindset, why is it people on here think humans have some kind of magical ability machines don't or can't? Five years ago I would never have predicted this kind of human chauvinism here. It's some kind of weird romanticism almost.

Maybe because everything LLM-written is written in the same style with no creativity, diversity, or idiosyncrasies? If all humans suddenly started writing in a single, bland, corporate style, that would be a tragedy, LLMs or not.

Re: Project Glasswing: what Mythos showed us

#90

What does this mean? > It's a different kind of tool doing a different kind of work, and that makes a clean apples-to-apples comparison to earlier models difficult. They claim it’s a different kind of tool and then describe using it the same way you’d use any other model. This really felt way worse than the average Cloudflare blog and really just rehashed the Mythos announcement which had already called out the key p…

My guess is because it is a model trained specifically for security/hacking. So comparing it to Opus, trained for chat/code/etc., is apples-to-oranges.

It is not, that's what surprised Anthropic employees too.
Post reply on HN