Live data from Hacker News

Project Glasswing: An Initial Update

anthropic.com

271–280 of 345 posts

Re: Project Glasswing: An Initial Update

#271
post #66

Earlier quoted context omitted.

I thought we were all doing that already?

Jesus, dude. There are managers reading this.

What do you think they do all day?

The larger pattern is not unique to writing code. Think of it next time a reorg comes, or some random thing gets "improved" in the name of "efficiency" only management seems to see.

Re: Project Glasswing: An Initial Update

#273
post #34

You can get a taste of this today yourself with Codex Security. I turned it on just as an experiment and in less than a week it has now become essential to all of us. I was shocked how accurate it is, how many security issues it found in existing code, how it continually finds them as we commit, and how NO ONE is immune from making these mistakes. I'd say it is about 90% accurate for us. Often even the "Low" findings…

> I expect tools like this to be a regular part of the development lifecycle from here on. We code with AI, we review with AI, we search for vulns with AI. Even if it isn't perfect, it is easily worth the cost IMHO. So, how is that supposed to work? Claude Code generates security bugs, then Claude Security finds them, then Claude Code generate fix, spend tokens, profit?

I wonder how many minivans Anthropic is going to code themselves.

Re: Project Glasswing: An Initial Update

#275
post #34

You can get a taste of this today yourself with Codex Security. I turned it on just as an experiment and in less than a week it has now become essential to all of us. I was shocked how accurate it is, how many security issues it found in existing code, how it continually finds them as we commit, and how NO ONE is immune from making these mistakes. I'd say it is about 90% accurate for us. Often even the "Low" findings…

> I expect tools like this to be a regular part of the development lifecycle from here on. We code with AI, we review with AI, we search for vulns with AI. Even if it isn't perfect, it is easily worth the cost IMHO. So, how is that supposed to work? Claude Code generates security bugs, then Claude Security finds them, then Claude Code generate fix, spend tokens, profit?

Software engineers generate security bugs, Software engineers find them, then Software engineers generate fix, collect salary, profit?

Re: Project Glasswing: An Initial Update

#276
post #147
post #142

Earlier quoted context omitted.

You’re right, it’s a valid data point. But the U.K. government report is also a data point, and the Firefox report is a data point, and they suggest that it is, indeed, significantly better than current generation models. Maybe curl is significantly better hardened than most projects? In any event, it barely matters. As Anthropic acknowledges, next level models are comings, theirs is only one of them. Current generat…

> Maybe curl is significantly better hardened than most projects? Meanwhile from [1]: "Not even half-way through this #curl release cycle we are already at 11 confirmed vulnerabilities - and there are three left in the queue to assess and new reports keep arriving at a pace of more than one/day." "The simple reason is: the (AI powered) tools are this good now. And people use these tools against curl source code.They…

Based on the article here, and Firefox's mythos article, they had found bugs with Opus 4.6 as well but mythos is finding more that it missed.

That would align with the curl feedback you linked, they aren't using mythos but are finding bugs with other models. Presumably the expectation would be that with mythos they'd find more that were missed by other models already used.

Re: Project Glasswing: An Initial Update

#277
The math doesn't add up. They say they found more than 20k vulnerabilities, then it decreases to 1700 high or critical, then this number becomes 175 (when Claude didn't reassess the CVE severity) and over 500 later on. Then they say they confirmed 800 vulnerabilities... what happened to the 20k figure?

Plus, they also mention they check if fixes are available for the bugs they found. What are the chances they are re-reporting old bugs just to inflate their numbers? Bugs that were already fixed?

And how can we be sure their reassessment is not artifically increasing the severity of the CVEs found just to create FUD and sell their product?

Re: Project Glasswing: An Initial Update

#278
post #274

> Progress on software security used to be limited by how quickly we could find new vulnerabilities And so was malicious vulnerability research.

the asymmetry stays the same though — defenders must find everything, attackers need one. LLMs accelerate both sides equally but that gap doesnt close

Re: Project Glasswing: An Initial Update

#279
Is there a reason why they appear to conflate vulnerabilities and bugs? It's not clear where they are defining their terms, eg

"After one month, most partners have each found hundreds of critical- or high-severity vulnerabilities in their software. Collectively, they’ve found more than ten thousand. Several have told us that their rate of bug-finding has increased by more than a factor of ten. For instance, Cloudflare has found 2,000 bugs (400 of which are high- or critical-severity) across their critical-path systems, with a false positive rate that Cloudflare’s team considers better than human testers." (emphasis mine)

Re: Project Glasswing: An Initial Update

#280
post #122

I’m not sure how to reconcile anthropic’s update / some of the exuberant comments here with recent feedback like the following from curl maintainer Daniel Steinberg: “I see no evidence that this setup [Mythos] finds issues to any particular higher or more advanced degree than the other tools have done before Mythos. Maybe this model is a little bit better, but even if it is, it is not better to a degree that seems to…

Fortunately, this is just a press release for their new product 'Claude Security'. Just contact sales to find out more https://claude.com/product/claude-security
Post reply on HN