OpenAI now tries to hide that ChatGPT was trained on copyrighted books
71–80 of 94 posts
Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books
#72Earlier quoted context omitted.
I'm still not understanding your argument here. If a ML model spits out snippets of copyrighted material and you try to monetize those then that's clearly infringement. But if it ingests copyrighted material and then spits out entirely new content influenced by the originals... how is that an issue?
> If a ML model spits out snippets of copyrighted material and you try to monetize those then that’s clearly infringement. If it spits it out because it was in the training set or prompt, yes. If it spits it out and it was in the training set (or prompt), it is at least difficult to make the case that it is not infringement (if it was the prompt, then you basically have the same problem as any other case where you ar…
1) Is ingestion of copyrighted material as part of the model training process an infringement itself? I say no - it is equivalent to a human going to a library and reading all the books there. The knowledge gained by the human is equivalent to the weights that end up in the model.
2) Is the output of copyrighted material from a model infringement? Yes, obviously. In the exact same way as a human regurgitating a duplicate or close-to-duplicate of someone else's book/song/script/whatever is. We already have an entire body of copyright law to cover this and don't need to reinvent the wheel just because we swapped a human creator for a machine.
Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books
#73Earlier quoted context omitted.
No they aren't. Where's the Terms of Service for a book? What about secondhand books or books you get from the library or books you borrow from a friend? I never had a book tell me to an accept a license agreement before I could read it.
Are you trolling? Look at the ISBN/info page of literally any book. It will say something like 'all rights reserved' and that you can't reproduce any part without permission, except for limited cases. Please go get a book off of your shelf and look, I'm begging you. Edit - Example from Infinite Jest: https://burnsiderarebooks.cdn.bibliopolis.com/pictures/14094...
Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books
#74Earlier quoted context omitted.
why computer in openAI have magic can let author lose their copyright ? you know, because they have a magical algorithm?
Yes. Deep learning is not copy and paste at all, and this is the whole point. If you ask ChatGPT to quote stuff from a famous book, of course you will get what you want. It is prompt engineering. Human are able to quote stuff verbatim too. But just like you cut open the brain, you open up the Numpy array weight, all you can see is just nonsense, you don't see any bookshelf sitting in there with a pile of copyrighted…
Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books
#75it would seem to me that from a technical perspective the weights of an AI model trained on copyrighted material would be a reproduction of the copyrighted work. just because you combine the information from millions (or more) copyrighted works together doesn't mean that you aren't reproducing them. the process of training requires reproduction and distribution of the works internally as part of the data processing p…
If I publish an article on a subject I extensively read about on books and add no new information, I’m just reproducing them. Should that be considered a violation of copyright?
Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books
#76People keep comparing it to a human ingesting content throughout their life and then being influenced in their own works. I'm sorry but that is not the same thing - at all. The concept of training a model with the explicit intent of selling the output of that model is inherently different. Not saying that it should be illegal. But it is clearly in violation of the spirit of existing copyright law, in my opinion. They…
Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books
#77People keep comparing it to a human ingesting content throughout their life and then being influenced in their own works. I'm sorry but that is not the same thing - at all. The concept of training a model with the explicit intent of selling the output of that model is inherently different. Not saying that it should be illegal. But it is clearly in violation of the spirit of existing copyright law, in my opinion. They…
> The concept of training a model with the explicit intent of selling the output of that model is inherently different Uh-oh, better tell the colleges and universities to stop promoting degree programs off the back of the potential increase in lifetime earnings. Wouldn't want anyone to get the idea that it's okay to absorb a bunch of copyrighted material during college to update their mental model, then make money by…
Let alone the huge shift of power from one to many (one university to many students), to many to one (many resources to one company).
Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books
#78Earlier quoted context omitted.
> If a ML model spits out snippets of copyrighted material and you try to monetize those then that’s clearly infringement. If it spits it out because it was in the training set or prompt, yes. If it spits it out and it was in the training set (or prompt), it is at least difficult to make the case that it is not infringement (if it was the prompt, then you basically have the same problem as any other case where you ar…
You're needlessly complicating this and also changing what the parent comment I was responding do takes issue with. There are two things to be evaluated here: 1) Is ingestion of copyrighted material as part of the model training process an infringement itself? I say no - it is equivalent to a human going to a library and reading all the books there. The knowledge gained by the human is equivalent to the weights that…
It might be Fair Use for other reasons, but any argument that uses an “it’s like a human doing X” analogy is, legally, misguided. The courts simply do not see machines as being legally analogous to human brains. Humans can “copy" content into their brains through their senses and its not only not a copyright violation, its not legally a copy for which you need to do anything like fair use analysis, copy it, even lossily, into a machine where it is non-transiently stored, then it is, legally, a copy, and a violation unless an exception like Fair Use applies.
> Is the output of copyrighted material from a model infringement? Yes, obviously.
Again, only if it is actually a copy and not coincidence, and only if exceptions like Fair Use don't apply. It may be that training is more likely to be Fair Use than inference is—I’ve seen good arguments that don’t rely on treating machines as analogs of humans for that—but if you are copying copyright protected material, lossily or not, with a machine, it's going to be infringement unless an exception to copyright, like Fair Use, applies.
Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books
#79Earlier quoted context omitted.
Did you, through copyright infringement, acquire your copies of Harry Potter? Did you use Harry Potter source material and a machine to transform the input and produce your parody material? I think it's different if OpenAI acquires licenses for all of the material it uses for training.
I sure didn't buy a license for every book and screenplay I've ever read. And I borrow a lot of books from the library. I'm not understanding why you think ML models need to license all of their input content when humans clearly do not.
ML training doesn't work without having a copy of the data (books and screenplays). That data can either be copied in a non-infringing way (buy the books; acquire a license) or in an infringing way (download the books from a corpus without explicit permission like these AI companies are accused of doing). LLM services need a license just like Facebook, Instagram, etc. need a license to republish the stuff you post. The copyright holder maintains their copyright and the service that republishes their work (Facebook, Instagram, etc.) or publishes derivative works (AI) needs a license.
Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books
#80Earlier quoted context omitted.
Could you explain the five basic plots? While the number of plot structures are fairly countable, a plot can still be considered different by changing part of its contents (character, setting, etc.). The decision on whether a plot outright infringes on an existing plot varies case-by-case. An example of outright copying is "Fistful of Dollars" directed by Sergio Leone, which lifts the plot from "Yojinbo" by Akira Kur…
1. David vs. Goliath 2. Romeo and Juliet 3. Robin Hood 4. Crime and Punishment 5. maybe there are only four.