Live data from Hacker News

LLM Usage in Debian: Three Proposals

debian.org

61–70 of 228 posts

Re: LLM Usage in Debian: Three Proposals

#61

> A LLM (...) merely produces syntactically likely combinations of the training data While doesn't matter much in the rest of the policy, this is a common misconception among AI skeptics. It is not the case for a long time (since RL is used heavily in the training) and a LLM may go beyond its training data.

I assume RL stands for reinforcement learning, which just means that the weights are adjusted by evaluating some loss function. If the loss function uses the training data (i.e. supervised learning) then the original assumption still holds. So I fail to see how this is a misconception.

I think OP might be referring to synthetic data, which is now used to provide more data for training?

Re: LLM Usage in Debian: Three Proposals

#62
post #23

This set of proposals, are (sorry) just stupid. It's like saying to someone, you are not allowed to saw wood using an electric saw, you must do it by hand. What are we doing here?! LLM(s) are just a tool. Use it as such. You should own the work anyway.

It's more like: food processing tool A is known to be defective. We ask employees of this restaurant and supply chains to avoid food processing tool A. This seems eminently reasonable. The question is whether the tool is in fact defective.

That's just it though. The tool(s) aren't necessarily defective. Just because you can chop off your own hand with an chainsaw or do a botch job of sawing a plank of wood because you weren't watching what you were doing with the electric saw, does not by itself make the tools (Chainsaw, Electric Saw) defective.

Re: LLM Usage in Debian: Three Proposals

#63
post #18

Don't misinterpret this link as representing a final decision. It's actually three separate proposals which will be debated and then voted on. Proposal A is "expressly forbid any contributions to Debian written with the use or assistance of large language models (LLMs) or other generative AI tools." Proposal B is "The Debian project allows AI-assisted contributions (partially or fully generated by an LLM), provided t…

The title says "proposals"

Why would someone misinterpret this as a final decision?

Re: LLM Usage in Debian: Three Proposals

#64

> A LLM (...) merely produces syntactically likely combinations of the training data While doesn't matter much in the rest of the policy, this is a common misconception among AI skeptics. It is not the case for a long time (since RL is used heavily in the training) and a LLM may go beyond its training data.

I assume RL stands for reinforcement learning, which just means that the weights are adjusted by evaluating some loss function. If the loss function uses the training data (i.e. supervised learning) then the original assumption still holds. So I fail to see how this is a misconception.

Supervised learning and reinforcement learning are mathematically distinct. Supervised pre-training as you said maximizes the likelihood of a token sequence. RL optimizes a policy against an external reward signal or execution environment. For example, RL evaluates code against runtime execution/unit tests/proof checkers (e.g lean). Models learn new strategies and are able to produce novel code if trained in an RL environment.

Re: LLM Usage in Debian: Three Proposals

#65

This set of proposals, are (sorry) just stupid. It's like saying to someone, you are not allowed to saw wood using an electric saw, you must do it by hand. What are we doing here?! LLM(s) are just a tool. Use it as such. You should own the work anyway.

Yes, they are. But that’s exactly what Proposal B (the only one without moral panic) says: allow its use, but place the responsibility on whoever uploads it. If you upload broken or unlicensed code, that’s on you. Proposal B is the most pragmatic and realistic option.

Re: LLM Usage in Debian: Three Proposals

#66

This set of proposals, are (sorry) just stupid. It's like saying to someone, you are not allowed to saw wood using an electric saw, you must do it by hand. What are we doing here?! LLM(s) are just a tool. Use it as such. You should own the work anyway.

> This set of proposals, are (sorry) just stupid. It's like saying to someone, you are not allowed to saw wood using an electric saw, you must do it by hand. Stupid? Since when has there been a lack of clarity in copyright status with electric vs. hand saws?

This has nothing to do with copyright status or copyrighted works. If it did, Debian wouldn't even be a thing. You can't be a puritan about this; How much of Debian can one say, hand on heart, is truly not inspired, or influenced by other people's work, copyrighted or not.

Re: LLM Usage in Debian: Three Proposals

#67

This set of proposals, are (sorry) just stupid. It's like saying to someone, you are not allowed to saw wood using an electric saw, you must do it by hand. What are we doing here?! LLM(s) are just a tool. Use it as such. You should own the work anyway.

Yes, they are. But that’s exactly what Proposal B (the only one without moral panic) says: allow its use, but place the responsibility on whoever uploads it. If you upload broken or unlicensed code, that’s on you. Proposal B is the most pragmatic and realistic option.

Yes. this I agree with. It's teh same thing many corporations like the one I work for are doing as well. Expecting and mandating that even if you use these so-called "AI" tools. that you are a) responsible and b) accountable for the "work".

Re: LLM Usage in Debian: Three Proposals

#68
post #27
post #18

Don't misinterpret this link as representing a final decision. It's actually three separate proposals which will be debated and then voted on. Proposal A is "expressly forbid any contributions to Debian written with the use or assistance of large language models (LLMs) or other generative AI tools." Proposal B is "The Debian project allows AI-assisted contributions (partially or fully generated by an LLM), provided t…

Of the 3 proposals, the 2nd/B seems to be more of an “informed consent” model while the other two are seeking a comprehensive exclusion. I hope earnest dialogue and driving to a broad consensus among significant contributors is forthcoming. Notable are the endorsements, which are balanced over all three options, perhaps indicating only 1/3 are supportive of a more permissive position. If this represents core contribu…

I wonder how long software projects will even be able to ban LLM use in an age where exploits are found by LLMs. Like, if you have two forks of Debian and one uses LLMs to fix exploitable bugs and the other one doesn't, the level of security they can offer will be worlds apart. And noone in their right mind would want to use the less secure one. Similar to how noone would want to drive a car that was 100% hand built by humans when we know that machines do a much better job at precision tasks. LLMs are just another tool in the end.

Re: LLM Usage in Debian: Three Proposals

#69
post #63
post #18

Don't misinterpret this link as representing a final decision. It's actually three separate proposals which will be debated and then voted on. Proposal A is "expressly forbid any contributions to Debian written with the use or assistance of large language models (LLMs) or other generative AI tools." Proposal B is "The Debian project allows AI-assisted contributions (partially or fully generated by an LLM), provided t…

The title says "proposals" Why would someone misinterpret this as a final decision?

Eternal September of the Spotless Mind.

Re: LLM Usage in Debian: Three Proposals

#70
post #63
post #18

Don't misinterpret this link as representing a final decision. It's actually three separate proposals which will be debated and then voted on. Proposal A is "expressly forbid any contributions to Debian written with the use or assistance of large language models (LLMs) or other generative AI tools." Proposal B is "The Debian project allows AI-assisted contributions (partially or fully generated by an LLM), provided t…

The title says "proposals" Why would someone misinterpret this as a final decision?

It seems the title wasn't always as clear.
Post reply on HN