Live data from Hacker News

Did Claude increase bugs in rsync?

alexispurslane.github.io

321–330 of 611 posts

Re: Did Claude increase bugs in rsync?

#321

Earlier quoted context omitted.

[flagged]

Is this comment LLM generated? Have fun with 1000x more Buns that literally no one is using or maintaining. An entire software industry built on top of a burning garbage pile of crappy, dead code.

> An entire software industry built on top of a burning garbage pile of crappy, dead code.

That has been the case for the last, oh, decade or so. Where do you think LLMs learned to slop code?

Re: Did Claude increase bugs in rsync?

#322

Earlier quoted context omitted.

You can use LLMs in multiple ways, from very hands on to make local changes to completely hands-off. I've seen plenty of code that was LLM generated but the commit message itself did not have the co-author attached to it. This only seems to happen when someone's interface to the codebase is completely though Claude/Codex/..., and those are usually the most verbose commits, and yet they say the least, because they jus…

I would expect a mature code base like rsync to have a lot of unit tests and integration tests and frankly if there's not enough that such bugs haven't been caught; that should be your first use of LLMs in order to setup some deterministic guidelines when you do start making changes to your actual code. I have been experimenting with both aforementioned styles with interesting results.

> I would expect a mature code base like rsync to have a lot of unit tests and integration tests

You might be surprised. C applications which interact heavily with the system - like rsync - can be tricky to test comprehensively, as it's nontrivial to inject faults into system calls. If the application is architected to support this kind of testing, or uses a HAL, that may make matters easier - but an older codebase like rsync probably isn't.

Re: Did Claude increase bugs in rsync?

#323

Earlier quoted context omitted.

Is this comment LLM generated? Have fun with 1000x more Buns that literally no one is using or maintaining. An entire software industry built on top of a burning garbage pile of crappy, dead code.

> An entire software industry built on top of a burning garbage pile of crappy, dead code. That has been the case for the last, oh, decade or so. Where do you think LLMs learned to slop code?

Things have been bad, but every company using its own bespoke LLM reimplementation of rsync and similar is so, so much worse.

Re: Did Claude increase bugs in rsync?

#324

If the author is this concerned about security, I’m curious why rsync doesn’t just build with fil-c by default and skip the noise. Those who need the extra perf to do more than 1 gigabit/s can build it in “unsafe” mode.

Because Fil-C is not a serious project

Re: Did Claude increase bugs in rsync?

#325

Earlier quoted context omitted.

Also from the article: > "Claude clearly made things worse" &emdash; the main claim This article was clearly generated by AI, yet I found no mention/attribution of that by author. How likely is it than someone who vibe codes articles would also vibe code the underlying analysis and be eager to accept an outcome that is highly validating of that person’s workflow? I’d say very.

Are the numbers wrong? That's the only relevant thing here. Also, humans do use em dashes, just FYI.

Yes, I do for example.

And the author discussed the use of AI pretty exhaustively in point 0 of the post.

Re: Did Claude increase bugs in rsync?

#326

Earlier quoted context omitted.

Agree. From the article: > Here's my favorite part, though. Digging into the data, one of the first things that jumped out at me with blinding clarity was that the worst release, by far, in rsync history was entirely prior to the introduction of Claude ... And yet nobody noticed. Language really does suggest the article's author does have a dog in this fight and is cloaking opinion in fancy statistics jargon. "Blindi…

Also from the article: > "Claude clearly made things worse" &emdash; the main claim This article was clearly generated by AI, yet I found no mention/attribution of that by author. How likely is it than someone who vibe codes articles would also vibe code the underlying analysis and be eager to accept an outcome that is highly validating of that person’s workflow? I’d say very.

He did admit as much:

> "The scripts used to fetch the data, collate it into a DuckDB database file, construct the views on that DB, and then do the statistical analysis on that data, were indeed written by GLM 5.1, as was the HTML and much of the original prose for the final report webpage you're looking at right now."

Re: Did Claude increase bugs in rsync?

#327

Earlier quoted context omitted.

Also the amount of commits is suspicious. In the last two months, rsync had about as much commits as in the last two years before that. Most of them written with claude. And then stuff like this is in there. That's exactly what I'd expect when someone is excited about AI usage and becomes... well, sloppy.

Tridge already explains this: "Like many developers of open source packages I’ve been hit by a flood of security reports lately in my role as the rsync maintainer. Many of those reports are AI generated (not all though, there are some notable ones with very careful and high quality manual analysis). As this flood started to get more intense I realised I needed to raise the defences on rsync a lot — we needed much mor…

I think Tridge is simultaneously trying to be proactive and kinda giving too much credit to marketing. Anthropic has not been able to really give numbers or actual values on what Mythos can really do. It just waved Mythos in front of the public like a boogeyman screaming that AI is going to cause a security nightmare (and it has, but mostly through vibe coded trash from what I’ve noticed); I’m hard pressed to find their statement that they spent less than $20,000 to find a Kerberos bug in FreeBSD a compelling win without a lot more context and they seem disinclined to provide that data. I really do wonder what evidence they have provided to their approved partners, all of this smells…weird.

I honestly think the main problem is Tridge just failed at communicating any of this correctly and I don’t think the implication he gives that all of this was due to the urgency of the impending security apocalypse really holds water.

Why was all of this written straight to the master branch? Now that the release is out, why not better explain what the urgency of this release was? Why wasn’t he proactive in communicating this and instead let the mob make up their own story? I think a lot of people are inclined to give Tridge a lot of leeway due to the fact that he literally is the reason why rsync exists, but this was avoidable and I think the comment in his response post where he mentions that, “I’d rather be out sailing than working on rsync security issues, so I have reached for several AI tools to help with what needs to be done,” speaks volumes as to what is going on.

Re: Did Claude increase bugs in rsync?

#328
post #305

Earlier quoted context omitted.

Perhaps in an “everyday language” way, but not in the technical, statistical sense. In an underpowered statistical study, a claim that two experimental conditions did not differ are not persuasive.

No. It's a description of the result of the maybe underpowered study. the underpowered study did not find evidence. Evidence is absent. Because it is underpowered, it's not evidence that the effect is absent. The claim is not "two experimental conditions did not differ". The claim is "The data do not show evidence that the experimental conditions did differ".

You say "the underpowered study did not find evidence". Not true, it found quite a bit of evidence - many statistics were presented. There is no absence of evidence. The author wrote about the evidence, presenting P values and other statistics.

Of course the critical part is not the numbers, but what they mean.

So, what does the evidence mean?

The author interprets it to mean that there is no difference. They state this several times:

"46% EXACT PERMUTATION TEST P-VALUE (ONE-SIDED, H₁: CLAUDE MEAN > HISTORICAL)[...] What this p-value tells us is There's nothing unusual about the Claude group."

"74% ONE-SIDED P-VALUE (H₁: CLAUDE MORE LIKELY ABOVE MEDIAN) Fisher's exact test asks: if we split all releases at the historical median (0.74 sev/10c), are these Claude releases significantly buggy than previous releases (more likely to land above the median)? With a p-value of 74%, the answer is a decisive no. "

In an under-powered study, when a P value is above your alpha level cutoff (.05, .01, whatever was chosen) you can't distinguish between "no effect" and "could be an effect, but I didn't see one".

Re: Did Claude increase bugs in rsync?

#330
post #256

I've been coding for over 2 decades. I love it, I've always loved it and I likely always will. I was an AI skeptic some months ago but truly Claude and Codex have changed my development style and velocity in a way I never imagined would ever be possible. With that, yes, I produce more code and am finding more bugs. So looking over at comments in HN articles the amount of polarising hate to anything produced with AI i…

I was similarly an AI skeptic 3 years ago. When GPT-4 was the state of the art, I thought we're going to plateau soon because of context size limits (remember back when you had to pay insane money just to get 32K)?

Last year was the first time I saw an AI agent actually debug and fix a non-trivial bug in a satisfactory way. Even then, trying to use it on larger tasks made it clear that it wasn't something I could just hand over the issue tracker to.

Now? I've been using Codex for the past several months to work on a nontrivial project. Which was prototyped in C++ (for library reasons mostly), then had the initial version written in Haskell, and more recently I got it ported to Rust to keep memory use in check on mobile.

These things are not trouble-free, but the sheer amount of progress made in just the last year alone is astounding. Skepticism is well and good, but healthy skepticism ought to yield to tangible evidence.

Post reply on HN