Earlier quoted context omitted.
I'm not a lawyer, but it seems like you stood up a straw man there. >Just because it's represented as a bunch of numbers does not make it non copyrightable. Can you give an example of where the bunch of numbers is copyrightable when it's not just a numeric encoding of something that was already copyrightable? Taking music and encoding it as a wav file is not a creative work, but it's a representation of a copyrighted…
The key factor of Feist v Rural is whether there was any original or creative process in the way the facts were arranged. Here, there's a whole lot of creative decisions in labelling and guiding of training that produces the weights, so it's reasonable to think it might be copyrightable. That is, the numbers are a whole lot more original than the issuance of phone numbers or part numbers.
AI weights are not open “source”
201–210 of 274 posts
Re: AI weights are not open “source”
#202The complexity described seems to be resting on the unestablished idea that weights are copyrightable in the first place. If they're not, then presumably "available weights", "ethical weights", and "open weights" are all the same: open weights. Either your weights are under NDA and presumably considered to be a trade secret, or they are public, and the words in your "license" mean absolutely nothing? That seems like…
What happens if we train a neural network on a single, copyrighted work? Say it has one input node (or even zero, if you like), and regardless of this input, its output is always exactly the copyrighted work it was trained on. What do its weights represent? Clearly, its weights represent a direct encoding of the original work. Those weights are copyrightable, but not by the person who trained the neural network -- the copyright is held by the owner of the original work.
What if we train the neural network on just two copyrighted works? If its one input node is 0, it outputs the first, and if it's 1, it outputs the 2nd. Almost certainly, its weights are a complicated, tangled mix encoding both, like a compression algorithm that completely rearranged its input. Who owns the copyright to those weights? To whatever extent the weights can be "factored out" into a set representing the first work and a set representing the second, clearly the copyright holder of the first work holds the copyright on the first "factored set", and the 2nd on the 2nd. It seems obvious that we must be able to do this "factoring out" somehow (even if the topology of the factored networks is different), because we know both works are exactly represented by the weights, and the neural network itself can use this information to reconstruct them both, so they're in there ... somewhere. So is there a sort of "joint copyright" on the combined weights, where nobody is really allowed to do anything with it without approval of the other? Regardless, it's still clear that whoever trained the neural network has no claim on any copyright.
Where is the breaking point extending this from 2 works to a billion? People make arguments like "drawing a car from memory isn't infringing on copyright design of that car", which ... are you sure? Reproducing a piece of music from memory (and selling it) is usually copyright infringement. You're allowed to learn a Taylor Swift song as part of your musical training, but you're not usually allowed to then play it back from memory and sell that recording (I'm not sure I morally agree with this treatment of covers, nor if it's globally applicable). So the argument that "surely neural networks are allowed to learn from copyrighted works" misses the point: they can learn all they want, but as soon as they reproduce verbatim (or close enough) a copyrighted work, they're infringing. And if they're representing a complete copy of the work within their weights (which they obviously are if they can reproduce it), then the original copyright holder has a claim on those weights. And never in this process has the trainer of the NN acquired any copyright to anything. The real trainer is a bunch of GPUs, after all.
If the neural network cannot reproduce any of the copyrighted works verbatim, then we're getting closer to "fair use" territory. Yes, it's permissible to write a summary of a copyrighted work. That is so lossy as to not "compete" with the original work in any meaningful way. If it could be demonstrated that neural networks do not encode completed works (no matter how hard the factorization would be), then one could make this argument. Unfortunately, the evidence is that LLMs are more than happy to completely regurgitate copyrighted works verbatim. It seems to me the copyright holder of the original work therefore must hold a share of the claim on the weights. Still, the GPUs that trained the network do not magically acquire copyright over anything.
I wonder if the real answer is that the weights are copyrighted, and that copyright is held jointly by hundreds of millions of people, and nobody can do anything with those weights without the approval of all the others. I'm not saying I like that universe, but I am saying it's the most internally consistent answer I can think of, and seems to follow from the above argument.
Re: AI weights are not open “source”
#203Earlier quoted context omitted.
I'm not a lawyer, but it seems like you stood up a straw man there. >Just because it's represented as a bunch of numbers does not make it non copyrightable. Can you give an example of where the bunch of numbers is copyrightable when it's not just a numeric encoding of something that was already copyrightable? Taking music and encoding it as a wav file is not a creative work, but it's a representation of a copyrighted…
The key factor of Feist v Rural is whether there was any original or creative process in the way the facts were arranged. Here, there's a whole lot of creative decisions in labelling and guiding of training that produces the weights, so it's reasonable to think it might be copyrightable. That is, the numbers are a whole lot more original than the issuance of phone numbers or part numbers.
Labeling training data may qualify for copyright, but if the underlying training data doesn’t taint the output as a derivative work then labeling isn’t going to qualify by itself.
Thus without some new and very generous interpretation AI companies are at best not going to benefit from copyright and at worst may be forced to create all training data in house. My suspicion is this generation of AI companies are in a very difficult situation.
Re: AI weights are not open “source”
#204Earlier quoted context omitted.
It seems very difficult to ensure that a model will never output any of the copyrighted content that it was trained on. I can only think of three ways, but perhaps there are others 1. Evaluate every output from the model to ensure that none of the outputs are copyrighted 2. Evaluate every input to a model to ensure that the inputs are either not copyrighted or properly licensed 3. Change the definition of copyright s…
Ensuring a model never outputs copyrighted content is unimportant and tangential. It's irrelevant. You don't look for a way to make humans output no copyrighted content, you address each time they do case by case. A model training being rendered fair use doesn't mean any of its output can be used for whatever regardless.
That's what I listed as #1 - evaluate each individual output of the model to see if it violates copyright.
Re: AI weights are not open “source”
#205Earlier quoted context omitted.
Compiled object code of a bunch of code you didn't write . I don't know why programmers are so eager to forget that copyright is not at all about what something is , and all about where it came from . It'd be hard to assert that you hold copyright over object code compiled from code you didn't write!
Or is the fact that compiled code enjoys copyright protection, even though it is not human generated, evidence that being generated by a human is not overly important for copyright protection?
Re: AI weights are not open “source”
#206The complexity described seems to be resting on the unestablished idea that weights are copyrightable in the first place. If they're not, then presumably "available weights", "ethical weights", and "open weights" are all the same: open weights. Either your weights are under NDA and presumably considered to be a trade secret, or they are public, and the words in your "license" mean absolutely nothing? That seems like…
> The complexity described seems to be resting on the unestablished idea that weights are copyrightable in the first place. Yes. Weights probably aren't copyrightable in the US. See Feist vs. Rural Telephone, in which the Supreme Court ruled that telephone directories are not copyrightable. The copyright clause in the Constitution ("To promote the Progress of Science and useful Arts, by securing for limited Times to…
Let me put a straw man, and try to find a middle point, when the copyright argument stops being applicable:
1. A painting was done by an artist.
2. On a computer.
3. With a help from an image processor software.
4. Using some advanced filters, like super-resolution, that utilize computer vision techniques. Like neural networks.
Many smartphones already automatically process your* photos with some advanced CV algorithms. That can be called "machine generated art".
I'd personally prefer to stop saying "neural network did X", same way as we don't say "a bulldozer built a road, a crane built a house".
Re: AI weights are not open “source”
#207I don't even think we should be using the word "ethical" because it implies that anything more permissive is unethical. We should call these morality clause licenses.
The question of whether or not we should have morality clauses involved is complicated. Most bad actors do not give a shit about the licensing status of the code they are using. And these licenses also cause headaches for people who want to follow the rules[0] and avoid copyleft trolling[1]. On the other hand, the morality clauses in OpenRAIL-M are relatively straightforward and non-obnoxious.
[0] This also applies to "non-commercial" licensing, since that is a concept entirely foreign to copyright law. As far as I'm concerned the 'NC' clause in Creative Commons just means 'OK to torrent'.
[1] A practice in which people abuse copyleft licenses to try and extract licensing agreements for minor license violations. The forgiveness periods added to GPLv3 and later versions of Creative Commons are specifically to prevent this behavior.
Re: AI weights are not open “source”
#208Earlier quoted context omitted.
The key factor of Feist v Rural is whether there was any original or creative process in the way the facts were arranged. Here, there's a whole lot of creative decisions in labelling and guiding of training that produces the weights, so it's reasonable to think it might be copyrightable. That is, the numbers are a whole lot more original than the issuance of phone numbers or part numbers.
The requirement for expertise doesn’t necessarily imply that that setting up perimeters for training AI is necessarily copyrightable. A normal brick wall for example needs skills to create but doesn’t qualify as the goal is not creative. If so the mechanical output of a process that doesn’t qualify for copyright is not going to qualify. Labeling training data may qualify for copyright, but if the underlying training…
It depends. If each individual training item has a small impact on the output coefficients, then perhaps it's not a derivative work of them. But if there's a large creative process in determining model training procedure, deciding labelling strategies, and applying those-- perhaps those numbers are strongly derived from those things.
Re: AI weights are not open “source”
#209The complexity described seems to be resting on the unestablished idea that weights are copyrightable in the first place. If they're not, then presumably "available weights", "ethical weights", and "open weights" are all the same: open weights. Either your weights are under NDA and presumably considered to be a trade secret, or they are public, and the words in your "license" mean absolutely nothing? That seems like…
Some thought experiments: What happens if we train a neural network on a single, copyrighted work? Say it has one input node (or even zero, if you like), and regardless of this input, its output is always exactly the copyrighted work it was trained on. What do its weights represent? Clearly, its weights represent a direct encoding of the original work. Those weights are copyrightable, but not by the person who traine…
On the other hand, we need the original input vector for this to work, and one could argue that the network weights are simply the algorithm for decoding the input vector into the copyrighted work. So the originator holds copyright on the input vector, not the weights. Does it matter if the input vector has smaller information content than the original work? Clearly this argument relies on the input vector being the "actual encoding", and therefore must have at least as much information. If the input vector is an embedding of "please show me the latest Tom Clancy novel in full", this argument breaks down.
Okay, this is hard.
Re: AI weights are not open “source”
#210Earlier quoted context omitted.
This is mostly right - It depends on what the weights represent and how they were generated so I would not go as far as the initial claim. A collection of numbers is copyrightable if it's the encoded result of a creative process. Just because it's represented as a bunch of numbers does not make it non copyrightable. That's why it says " original works of authorship fixed in any tangible medium of expression, now know…
I'm not a lawyer, but it seems like you stood up a straw man there. >Just because it's represented as a bunch of numbers does not make it non copyrightable. Can you give an example of where the bunch of numbers is copyrightable when it's not just a numeric encoding of something that was already copyrightable? Taking music and encoding it as a wav file is not a creative work, but it's a representation of a copyrighted…
Sure, there are "poems" that consist of just a groups of numbers that are copyrighted. They are not encodings, it's just a string of numbers. It's indistinguishable from a bunch of numbers. This is just one example, there are lots.
They are enforceable to the degree it's creative, and to the degree the infringing use is also creative.
So you would not be able to sue me for using those numbers in a math equation. You would be able to sue me for reproducing your poem in a book of poems :)
As feist says, the creativity required for copyright is quite minimal. But it's still only as protectable as it is creative.
Look - AI is not the first thing to have this "issue". The answer remains the same as it always was - it's mostly about the process not the output.
The output mostly matters is if the output is not intended to be creative (or it's de minimis or ...).
Copyright as it currently exists is weird.
Like if you go to the copyright office and try to register your ssh public key and say "this was generated by ssh-keygen i had nothing to do with it" you may get a different result than if you said "this is my new visually stunning masterpiece, my ssh public key, which was generated with computer help but I used 37 precisely timed keyboard smashes to do it. Prints are available from my gallery for $500"