Live data from Hacker News

Kernel code removals driven by LLM-created security reports

lwn.net

51–60 of 130 posts

Re: Kernel code removals driven by LLM-created security reports

#53

Earlier quoted context omitted.

My experience with these tools is that they generate absolutely enormous amounts of insidiously wrong false positives, and it actually takes a decent amount of skill to work through the 99% which is garbage with any velocity. Of course some people don't do that, and send all the reports anyway... and then scream from the hilltops about how incredible LLMs are when by sheer luck one happens to be right. Not only is th…

This is incorrect. Here's the curl maintainer talking about dozens of bugs found using LLMs: https://daniel.haxx.se/blog/2025/10/10/a-new-breed-of-analyz...

From the curl blog post:

> "Remarkably few of them complete false positives."

Re: Kernel code removals driven by LLM-created security reports

#54
post #52

When LLM reports the bug, is should be used to fix it on the same occasion. Nobody will bother afterwards.

LLM being able to find bug doesn't necessarily equate to LLM being able to satisfactorily fix bug. Be happy that the bugs are being uncovered in the first place and brought to the attention of those who are concerned with their resolution.

Re: Kernel code removals driven by LLM-created security reports

#55

Earlier quoted context omitted.

My experience with these tools is that they generate absolutely enormous amounts of insidiously wrong false positives, and it actually takes a decent amount of skill to work through the 99% which is garbage with any velocity. Of course some people don't do that, and send all the reports anyway... and then scream from the hilltops about how incredible LLMs are when by sheer luck one happens to be right. Not only is th…

Your experience seems to be at least 3-6 months old. Long time kernel maintainers have recently written on this subject. They say that ~3 months ago the quality and accuracy of the reports crossed a threshold and are now legitimately useful.

The experience I'm describing was two weeks ago.

Yes, what we see coming out of the bottom of funnel is now is a little better. But it's sort of like reading day trading blogs: nobody shares their negative results, which in my direct experience are so bad they almost negate any investigative benefit. I also think part of this is that a small set of very prolific spammers were sufficiently discouraged to stop.

Re: Kernel code removals driven by LLM-created security reports

#56
post #25

Earlier quoted context omitted.

The reason Mythos isn't being released publicly is to drive up Anthropic's valuation by making big promises.

https://blog.mozilla.org/en/privacy-security/ai-security-zer... > As part of our continued collaboration with Anthropic, we had the opportunity to apply an early version of Claude Mythos Preview to Firefox. This week’s release of Firefox 150 includes fixes for 271 vulnerabilities identified during this initial evaluation.

So you're saying Mozilla is in on it, hyping up Anthropic. Are they getting a kickback?

Re: Kernel code removals driven by LLM-created security reports

#58

Earlier quoted context omitted.

In terms of quantity, definitely yes (a single person managing a swarm of Opusi can already find much more real bugs than a security researcher, hence the rise in reports). In terms of quality ("are there bugs that professional humans can't see at any budget but LLMs can?") - it's not very clear, because Opus is still worse than a human specialist, but Mythos might be comparable. We'll just have to wait and see what…

> Opusi The plural of "Opus" is "Opera". Might be a tad confusing tho :)

Wondered for a second "what does that browser have to do with all this?"

Re: Kernel code removals driven by LLM-created security reports

#59

Earlier quoted context omitted.

This is incorrect. Here's the curl maintainer talking about dozens of bugs found using LLMs: https://daniel.haxx.se/blog/2025/10/10/a-new-breed-of-analyz...

From the curl blog post: > "Remarkably few of them complete false positives."

That's worse than a report that can be easily dismissed
Post reply on HN