Earlier quoted context omitted.
So what? I can probably produce parts of the header from memory. Doesn't mean my brain is GPLed.
> So what? I can probably produce parts of the header from memory. Doesn't mean my brain is GPLed. Your brain is part of you. Some might say it is your very essence. You are human. Humans have inalienable rights that sometimes trump those enshrined by copyright. One such right is the right to remember things you've read. LLMs are not human, and thus don't enjoy such rights. Moreover, your brain is not distributed to…
The current state of the theory that GPL propagates to AI models
291–300 of 314 posts
Re: The current state of the theory that GPL propagates to AI models
#292Earlier quoted context omitted.
Ideally, Congress would just settle this basket of copyright concerns, as they explicitly have the power to do—and have done so repeatedly in the specific context of computers and software.
I've pitched this idea before but my pie in the sky hope is to settle most of this with something like a huge rollback of copyright terms, to something like 10 or 15 years initially. You can get one doubling of that by submitting your work to an official "library of congress" data set which will be used to produce common, clean, and open models that are available to anyone for a nominal fee and prevent any copyright…
> You can get one doubling of that by submitting your work to an official "library of congress" data set
Needs to be done at day 0 and made available at day 0 for usage. Maybe with standardised availability for usage e.g. licencing X,Y or Z or non-standard call-us-for-pricing.
The world moves fast and time really matters e.g. look at how the wait for patents to expire affects outcomes
Re: The current state of the theory that GPL propagates to AI models
#293Earlier quoted context omitted.
Bad analogy, probably made up by capitalists to confuse people. ML models cannot and do not learn. "learning" is a name of a process, when model developer downloads pirated material and processes it with an algorithm (computes parameters from it). Also, humans do not need to read million of pirated books to learn to talk. And a human artist doesn't need to steal million pictures to learn to draw.
> And a human artist doesn't need to steal million pictures to learn to draw. They... do? Not just pictures, but also real life data, which is a lot more data than an average modern ML system has. An average artist has probably seen- stolen millions of pictures from their social media feeds over their lifetime. Also, claiming to be anti-capitalist while defending one of the most offensive types of private property th…
Owning a song, a book or a picture doesn't give you much power by itself.
Re: The current state of the theory that GPL propagates to AI models
#294Earlier quoted context omitted.
Even if they are not "like" human brains in some sense, are they "like" brains enough to be counted similarly in a legal environment? Can you articulate the difference as something other than meat parochialism, which strikes me as arbitrary?
If LLMs are like human minds enough, then legally speaking we are abusing thinking and feeling human-like beings possessing will and agency in ways radically worse than slavery. What is missing in the “if I can remember and recite program then they must be allowed to remember and recite proframs” argument is that you choose to do it (and you have basic human rights and freedoms), and they do not.
Re: The current state of the theory that GPL propagates to AI models
#295Earlier quoted context omitted.
Human learning is materially different from LLM training. They're similar in that both involve providing input to a system that can, afterwards, produce output sharing certain statistical regularities with the input, including rote recital in some cases – but the similarities end there.
>Human learning is materially different from LLM training [...] but the similarities end there. Specifically what "material differences" are there? The only arguments I heard are are around human exceptionalism (eg. "brains are different, because... they just are ok?"), or giving humans a pass because they're not evil corporations.
So, I'll provide an example: humans can learn to do mathematics. LLMs cannot. This example is particularly galling because there are computer programs that can do (some, limited) mathematics: those operate largely by brute-force, yet can solve more mathematics problems using fewer resources than LLMs.
Re: The current state of the theory that GPL propagates to AI models
#296Earlier quoted context omitted.
> A small shop has own cctv to catch intruders = one thing. Local company installing cctv everywhere = different thing. But that's the thing you were implying couldn't be distinguished. Every small shop having its own CCTV is different than one company having cameras everywhere, even if they both result cameras all over the place. > "Malware exists and nobody can unexist it now because it's just code and data" Which…
like LLM or NFT or killer drones, malware isn't bad for somebody . it is always about who it is benefits the most. > the LLMs that aren't bad which LLM is not made by stealing copyleft code?
If the incumbent copyright interests insist on picking an unnecessary fight with LLMs or AI in general, they will and must lose decisively. That applies to all of the incumbents, from FSF to Disney. Things are different now.
Re: The current state of the theory that GPL propagates to AI models
#297Earlier quoted context omitted.
> even if your model was trained strictly on copyleft material That's not legal use of the material according to most copyleft licenses. Regardless if you end up trying to reproduce it. It's also quite immoral if technically-strictly-speaking-maybe-not-unlawful.
> That's not legal use of the material according to most copyleft licenses. That probably doesn't matter given the current rulings that training an AI model on otherwise legally acquired material is "fair use", because the copyleft license inherently only has power because of copyright. I'm sure at some point we'll see litigation over a case where someone attempts to make "not using the material to train AI" a term o…
Re: The current state of the theory that GPL propagates to AI models
#298Earlier quoted context omitted.
> A small shop has own cctv to catch intruders = one thing. Local company installing cctv everywhere = different thing. But that's the thing you were implying couldn't be distinguished. Every small shop having its own CCTV is different than one company having cameras everywhere, even if they both result cameras all over the place. > "Malware exists and nobody can unexist it now because it's just code and data" Which…
like LLM or NFT or killer drones, malware isn't bad for somebody . it is always about who it is benefits the most. > the LLMs that aren't bad which LLM is not made by stealing copyleft code?
Malware isn't bad for Russian crime syndicates, but we're generally content to regard them as the adversary and not care about their satisfaction. That isn't the case for someone who wants to use an LLM to fix a bug in their printer. They're doing the good work and people trying to stop them are the adversary.
> which LLM is not made by stealing copyleft code?
Let's drive a stake through this one by going completely the other way. Suppose you train an LLM only on GPL code, and all the people distributing and using it are only distributing its output under the GPL. Regardless of whether that's required, it's allowed, right? How would you accuse any of those people of a GPL violation?
Re: The current state of the theory that GPL propagates to AI models
#299Earlier quoted context omitted.
Now I'm kind of curious if you give an LLM the disassembly of a proprietary firmware blob and tell it to turn it into human-readable source code, how good is it at that? You could probably even train one to do that in particular. Take existing open source code and its assembly representations as training data and then treat it like a language translation task. Use the context to guess what the variable names were bef…
The most difficult parts of getting readable code would be dealing with inlined functions and otherwise-duplicated code from macros or similar, and dealing with in-memory structure layouts; both pretty complicated very-global tasks. (never mind naming things, but perhaps LLMs have a good shot at that) That said, chatgpt currently seems to fail even basic things - completely missed the `thrM` path being possible here:…
https://chatgpt.com/s/t_6929f00ff5508191b75f31e219609a35 (5.1 Pro Thinking)
https://claude.ai/share/7d9caa25-14f7-4233-b15c-d32b86e20e09 (Opus 4.5)
https://docs.google.com/document/d/1C0lSKbLSZOyMWnGgR0QhZh3Q... (Gemini 3 Pro Thinking)
All of them recognized the thrM exception path, although I didn't review them for correctness.
That being said, I imagine the major showstopper in real-world disassembly tasks would simply be the limited context size. As you suggest, a standard LLM isn't really the best tool for the job, at least not without assistance to split up the task logically.
Re: The current state of the theory that GPL propagates to AI models
#300We need a new license that forbids all training. That is the only way to stop big corporations from doing this.
Fair use doesn’t need a license, so it doesn’t matter what you put in the license. Generally speaking licenses give rights (they literally grant license ). They can’t take rights away, only the legislature can do that.