Live data from Hacker News

LLM Usage in Debian: Three Proposals

debian.org

201–210 of 228 posts

Re: LLM Usage in Debian: Three Proposals

#201

Earlier quoted context omitted.

Some people want to uphold the rule even if they can't be caught.

Some people find a duty to ignore senseless rules. If Debian wants to make the LLM thing good, sue the providers for a billion or two for copyright violation, and join a class action to do some ASCAP style thing routing their money back to copyright holders. Separately, lobby to regulate data center build outs to include the solar and battery they need, and not be built in water stressed areas or near people who obje…

It's strange we're complaining about too much investment. The reason we avoid deflation is because it hinders investment, but perhaps it's time for a bit less money creation and have a period of deflation, even if it makes the "growth" figures look bad.

Politically difficult of course.

Re: LLM Usage in Debian: Three Proposals

#202

A lot of people pointing out that Proposal A is untenable/ at philosophical odds with the widespread upstream LLM usage, etc. etc. I think that's an unfortunate framing & I think A would be the best option for the opposite reasons. If they go with Proposal A, people will absolutely use AI tooling to contribute to Debian. This might be in bad faith (unlikely but possible), but it will also be done in good faith throug…

The wording of "written with ai tools" implies generation, as do all the concerns about copyright etc. it would be beyond insane to attempt to ban asking a chat bot or an agent questions to navigate or understand a codebase, not in the least because that is completely unenforceable

> not in the least because that is completely unenforceable

You've completely missed my point. It's ALL completely unenforceable. That doesn't mean the directive is useless - it's a guideline for contributors towards a general project direction that will materially limit & reduce overt LLM usage by anyone acting in good faith (the vast majority). The enforcement doesn't need to be strict, but anything less than strict wording won't be effective in influencing people's behaviour.

Re: LLM Usage in Debian: Three Proposals

#203
post #151

Earlier quoted context omitted.

> Notable are the endorsements, which are balanced over all three options... I don't think this is notable. The process requires five seconders, so you're just seeing the set of people who seconded asynchronously before there were obviously enough that others didn't bother. The culture is to avoid unnecessary noise and leave it for the vote.

A bit of topic, but does the word "seconded" still work when five are required? Isn't there a more generic term?

Depends on your rules of order. Under Robert's you can have multiple seconds but seconds aren't recorded. In model UN there are multiple seconds. Second'ing a motion is not generally an endorsement per se, it's just saying you support a vote on the matter.

Re: LLM Usage in Debian: Three Proposals

#204

Earlier quoted context omitted.

Some of the reasons to ban LLMs are beyond what they do to code. I.e. environmental degradation, scraping without authorization, reduce human interaction etc. So I think banning them still makes sense.

All of which have been challenged though. As the stake is high, any such reason (for or against LLMs, to be fair) has to be backed with Wikipedia-level evidences at the very least.

I disagree. We should adopt caution - so that the burden of proof is on those doing the inventing.

Re: LLM Usage in Debian: Three Proposals

#205

Earlier quoted context omitted.

I wonder how long software projects will even be able to ban LLM use in an age where exploits are found by LLMs. Like, if you have two forks of Debian and one uses LLMs to fix exploitable bugs and the other one doesn't, the level of security they can offer will be worlds apart. And noone in their right mind would want to use the less secure one. Similar to how noone would want to drive a car that was 100% hand built…

There are a few of these "maintainer said no AI pulls, so we forked it" in the wild already and I am really interested to see how it all pans out. Like, okay, sure, maintainer says AI bug monitoring is cumbersome. One fork stops reading the AI bug monitoring, the other leans in. Isn't the latter more likely to find/fix critical/performance/security issues? >Similar to how noone would want to drive a car that was 100%…

> There are a few of these "maintainer said no AI pulls, so we forked it" in the wild already and I am really interested to see how it all pans out.

There are also forks that go in the other direction, i.e. forking to avoid AI contributions upstream: https://drewdevault.com/blog/Forking-vim/

And promises to do so if it becomes necessary: https://human-emacs.org/

Will indeed be interesting to see how it all pans out.

Re: LLM Usage in Debian: Three Proposals

#206
post #186

Earlier quoted context omitted.

I simply don't accept that this type of project rule has any validity. It's like telling me what editor I can use, or that I can't use Google to search for programming tips, or that I can only use an electric car to drive to the office.

Debian has, like all software projects, always had rules about how you created the code you contributed. No software project is going to accept your contributed code which you wrote at work using company time and resources; that’s a copyright violation. There has never been a situation in which simply your contributed code can be analyzed and accepted in a vacuum. Bits do have color: https://ansuz.sooke.bc.ca/entry/2…

They are unenforceable and the motivation looks to be political in nature. I will simply ignore them, but it does reduce the credibility of Debian if the more AI-hostile variants of these pass.

Re: LLM Usage in Debian: Three Proposals

#207
post #117

Earlier quoted context omitted.

This is a somewhat common talking point for supporting LLMs. What did non-English speakers do before LLMs?

They learned English or gave up.

What about translation software that existed before LLMs?

Re: LLM Usage in Debian: Three Proposals

#208
post #168
post #118

From Proposal A: > LLM output has very unclear legal status: it may be possible to copyright on its own merits, or not; it may be affected by all of the licenses and copyrights in the training data, or not. Debian Policy and the DFSG require absolute clarity for licensing and copyright[1][2]. Software and other contributions written conventionally by humans with unclear copyright or license status are not allowed in…

At the same time, how can a human know if they accidentally rewrote exactly some code they read in the past from a codebase with an incompatible license? Accidental copies happen all the time in code, music and art. Not all of it is necessarily super original, but that's beside the point and it happens. So while we think we can hold AI tools to higher standards than humans, they are still modeled after human thinking…

This argument is almost algorithmic: Take the concern about , substitute 'human'. Also, another popular response: don't address questions, instead try to put the person asking on the defensive.

The main point is the implicit argument that humans and machines are equivalent.

Re: LLM Usage in Debian: Three Proposals

#210
post #190
post #116

Earlier quoted context omitted.

Doesn't the LLM still output the next most probable token (modulus heat/desired variation) according to its model? Training shifts the probability - but doesn't change the fact that the output is a sampling based on input and the model?

So do humans more or less? The only counterpoints I've heard are religious or unmeasured quantum brain something.

No. Humans actually understand things and can reason from one fact to another. We do not just spit out the next most likely thing without any intelligence the way an LLM does.
Post reply on HN