Live data from Hacker News

The current state of the theory that GPL propagates to AI models

shujisado.org

291–300 of 314 posts

Re: The current state of the theory that GPL propagates to AI models

#291
post #200

Earlier quoted context omitted.

So what? I can probably produce parts of the header from memory. Doesn't mean my brain is GPLed.

> So what? I can probably produce parts of the header from memory. Doesn't mean my brain is GPLed. Your brain is part of you. Some might say it is your very essence. You are human. Humans have inalienable rights that sometimes trump those enshrined by copyright. One such right is the right to remember things you've read. LLMs are not human, and thus don't enjoy such rights. Moreover, your brain is not distributed to…

How do you know I'm human?

Re: The current state of the theory that GPL propagates to AI models

#292

Earlier quoted context omitted.

Ideally, Congress would just settle this basket of copyright concerns, as they explicitly have the power to do—and have done so repeatedly in the specific context of computers and software.

I've pitched this idea before but my pie in the sky hope is to settle most of this with something like a huge rollback of copyright terms, to something like 10 or 15 years initially. You can get one doubling of that by submitting your work to an official "library of congress" data set which will be used to produce common, clean, and open models that are available to anyone for a nominal fee and prevent any copyright…

Fighting corporates seems like a loser move: the corporate winners (from long copyright terms) would deploy income to politically prevent disadvantageous change.

> You can get one doubling of that by submitting your work to an official "library of congress" data set

Needs to be done at day 0 and made available at day 0 for usage. Maybe with standardised availability for usage e.g. licencing X,Y or Z or non-standard call-us-for-pricing.

The world moves fast and time really matters e.g. look at how the wait for patents to expire affects outcomes

Re: The current state of the theory that GPL propagates to AI models

#293
post #114

Earlier quoted context omitted.

Bad analogy, probably made up by capitalists to confuse people. ML models cannot and do not learn. "learning" is a name of a process, when model developer downloads pirated material and processes it with an algorithm (computes parameters from it). Also, humans do not need to read million of pirated books to learn to talk. And a human artist doesn't need to steal million pictures to learn to draw.

> And a human artist doesn't need to steal million pictures to learn to draw. They... do? Not just pictures, but also real life data, which is a lot more data than an average modern ML system has. An average artist has probably seen- stolen millions of pictures from their social media feeds over their lifetime. Also, claiming to be anti-capitalist while defending one of the most offensive types of private property th…

For property to give you power, you need to concentrate lot of it, or own something expensive or exclusive. That's what capitalists do: a capitalist owns an expensive lathe, and you don't so you have no choice but to work for the capitalist on capitalist's terms (you can replace lathe with GPU farm for more modern analogy).

Owning a song, a book or a picture doesn't give you much power by itself.

Re: The current state of the theory that GPL propagates to AI models

#294

Earlier quoted context omitted.

Even if they are not "like" human brains in some sense, are they "like" brains enough to be counted similarly in a legal environment? Can you articulate the difference as something other than meat parochialism, which strikes me as arbitrary?

If LLMs are like human minds enough, then legally speaking we are abusing thinking and feeling human-like beings possessing will and agency in ways radically worse than slavery. What is missing in the “if I can remember and recite program then they must be allowed to remember and recite proframs” argument is that you choose to do it (and you have basic human rights and freedoms), and they do not.

We're halfway to Roko's Basilisk here

Re: The current state of the theory that GPL propagates to AI models

#295
post #180

Earlier quoted context omitted.

Human learning is materially different from LLM training. They're similar in that both involve providing input to a system that can, afterwards, produce output sharing certain statistical regularities with the input, including rote recital in some cases – but the similarities end there.

>Human learning is materially different from LLM training [...] but the similarities end there. Specifically what "material differences" are there? The only arguments I heard are are around human exceptionalism (eg. "brains are different, because... they just are ok?"), or giving humans a pass because they're not evil corporations.

We don't understand human brains well enough to answer this question specifically: if we understood the mechanisms of the human brain, we could replicate them in software, and AI would be that much more advanced. We know that real human neurons don't work at all like artificial neural networks, but that isn't a proof: phonons aren't bosons or fermions (their spin isn't well-defined), yet a sonic black hole is a useful model for Hawking radiation.

So, I'll provide an example: humans can learn to do mathematics. LLMs cannot. This example is particularly galling because there are computer programs that can do (some, limited) mathematics: those operate largely by brute-force, yet can solve more mathematics problems using fewer resources than LLMs.

Re: The current state of the theory that GPL propagates to AI models

#296

Earlier quoted context omitted.

> A small shop has own cctv to catch intruders = one thing. Local company installing cctv everywhere = different thing. But that's the thing you were implying couldn't be distinguished. Every small shop having its own CCTV is different than one company having cameras everywhere, even if they both result cameras all over the place. > "Malware exists and nobody can unexist it now because it's just code and data" Which…

like LLM or NFT or killer drones, malware isn't bad for somebody . it is always about who it is benefits the most. > the LLMs that aren't bad which LLM is not made by stealing copyleft code?

You don't get to unilaterally make laws for the rest of us, which is what you are trying to do when you throw around terms like "stealing" in contexts where they have no legal meaning. Sorry.

If the incumbent copyright interests insist on picking an unnecessary fight with LLMs or AI in general, they will and must lose decisively. That applies to all of the incumbents, from FSF to Disney. Things are different now.

Re: The current state of the theory that GPL propagates to AI models

#297
post #129

Earlier quoted context omitted.

> even if your model was trained strictly on copyleft material That's not legal use of the material according to most copyleft licenses. Regardless if you end up trying to reproduce it. It's also quite immoral if technically-strictly-speaking-maybe-not-unlawful.

> That's not legal use of the material according to most copyleft licenses. That probably doesn't matter given the current rulings that training an AI model on otherwise legally acquired material is "fair use", because the copyleft license inherently only has power because of copyright. I'm sure at some point we'll see litigation over a case where someone attempts to make "not using the material to train AI" a term o…

Indeed, the GPL's definitions of "modify" and "propagate" restrict the license's scope to actions that would otherwise infringe on copyright if not permitted. And fair use and similar doctrines generally act as carve-outs to copyright infringement.

Re: The current state of the theory that GPL propagates to AI models

#298

Earlier quoted context omitted.

> A small shop has own cctv to catch intruders = one thing. Local company installing cctv everywhere = different thing. But that's the thing you were implying couldn't be distinguished. Every small shop having its own CCTV is different than one company having cameras everywhere, even if they both result cameras all over the place. > "Malware exists and nobody can unexist it now because it's just code and data" Which…

like LLM or NFT or killer drones, malware isn't bad for somebody . it is always about who it is benefits the most. > the LLMs that aren't bad which LLM is not made by stealing copyleft code?

> like LLM or NFT or killer drones, malware isn't bad for somebody.

Malware isn't bad for Russian crime syndicates, but we're generally content to regard them as the adversary and not care about their satisfaction. That isn't the case for someone who wants to use an LLM to fix a bug in their printer. They're doing the good work and people trying to stop them are the adversary.

> which LLM is not made by stealing copyleft code?

Let's drive a stake through this one by going completely the other way. Suppose you train an LLM only on GPL code, and all the people distributing and using it are only distributing its output under the GPL. Regardless of whether that's required, it's allowed, right? How would you accuse any of those people of a GPL violation?

Re: The current state of the theory that GPL propagates to AI models

#299
post #273

Earlier quoted context omitted.

Now I'm kind of curious if you give an LLM the disassembly of a proprietary firmware blob and tell it to turn it into human-readable source code, how good is it at that? You could probably even train one to do that in particular. Take existing open source code and its assembly representations as training data and then treat it like a language translation task. Use the context to guess what the variable names were bef…

The most difficult parts of getting readable code would be dealing with inlined functions and otherwise-duplicated code from macros or similar, and dealing with in-memory structure layouts; both pretty complicated very-global tasks. (never mind naming things, but perhaps LLMs have a good shot at that) That said, chatgpt currently seems to fail even basic things - completely missed the `thrM` path being possible here:…

You can't feed something like that to the free ChatGPT model and expect anything useful. Try these:

https://chatgpt.com/s/t_6929f00ff5508191b75f31e219609a35 (5.1 Pro Thinking)

https://claude.ai/share/7d9caa25-14f7-4233-b15c-d32b86e20e09 (Opus 4.5)

https://docs.google.com/document/d/1C0lSKbLSZOyMWnGgR0QhZh3Q... (Gemini 3 Pro Thinking)

All of them recognized the thrM exception path, although I didn't review them for correctness.

That being said, I imagine the major showstopper in real-world disassembly tasks would simply be the limited context size. As you suggest, a standard LLM isn't really the best tool for the job, at least not without assistance to split up the task logically.

Re: The current state of the theory that GPL propagates to AI models

#300

We need a new license that forbids all training. That is the only way to stop big corporations from doing this.

Fair use doesn’t need a license, so it doesn’t matter what you put in the license. Generally speaking licenses give rights (they literally grant license ). They can’t take rights away, only the legislature can do that.

Exclusive or co-exclusive licences can nullify your default fair use in certain jurisdictions.
Post reply on HN