I was also trained on copyrighted books. What’s the problem? Isn’t the whole purpose of a book to be read?
Are you a bot?
OpenAI now tries to hide that ChatGPT was trained on copyrighted books
61–70 of 94 posts
Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books
#62Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books
#63Earlier quoted context omitted.
I'm not understanding how it's different. For example, if I specifically set out to make money by creating and selling a parody of Harry Potter by reading all the Harry Potter books a bunch of times, does that make my parody a violation of the spirit of existing copyright law? Edit to add another example because someone is going to say that parody is its own thing and exempted. If I want to make money by writing a fi…
Surely you must understand the difference between copying something and producing something entirely new. If you literally only consumed Tarantino media and had no other influences, and then one day made a movie so great everyone was calling you the next Tarantino - that would be fine! As long as you didn't copy anything that he actually made himself. Being influenced by is not the same as copying.
Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books
#64Earlier quoted context omitted.
Please explain how this is any different than the Google Books case. Please do so without any feelings involved to the best of your ability. Even if OpenAI maliciously ingested copywritten work, it is just a bunch of numbers and if you go in and ask for it spit the book back out, it won't. That simply isn't how this works.
I don't see much of a difference, seems like google profited off of copyrighted material that they didn't own. I dont understand the purpose of your feelings comment. I have no affiliations or preferences for any dogs in this fight. I don't necessarily think they are identical scenarios but if I were OpenAI's lawyer that's probably what I'd try and point at. I didnt say it would spit the book back out, at any point.…
OpenAI isn't pirating any books content nor is it distributing or reproducing it. Nor are they even "stealing the book".
Both are transformative content.
Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books
#65Earlier quoted context omitted.
I'm not understanding how it's different. For example, if I specifically set out to make money by creating and selling a parody of Harry Potter by reading all the Harry Potter books a bunch of times, does that make my parody a violation of the spirit of existing copyright law? Edit to add another example because someone is going to say that parody is its own thing and exempted. If I want to make money by writing a fi…
Did you, through copyright infringement, acquire your copies of Harry Potter? Did you use Harry Potter source material and a machine to transform the input and produce your parody material? I think it's different if OpenAI acquires licenses for all of the material it uses for training.
Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books
#66Earlier quoted context omitted.
Yes, it is different that human has magic and machine don't. I learn stuff from books and sell my skills. I definitely copied someone's knowledge about calculus, engineering, and I am intented to earn money from them. Machine cannot do this because only me has magic.
why computer in openAI have magic can let author lose their copyright ? you know, because they have a magical algorithm?
Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books
#67Earlier quoted context omitted.
Surely you must understand the difference between copying something and producing something entirely new. If you literally only consumed Tarantino media and had no other influences, and then one day made a movie so great everyone was calling you the next Tarantino - that would be fine! As long as you didn't copy anything that he actually made himself. Being influenced by is not the same as copying.
I'm still not understanding your argument here. If a ML model spits out snippets of copyrighted material and you try to monetize those then that's clearly infringement. But if it ingests copyrighted material and then spits out entirely new content influenced by the originals... how is that an issue?
If it spits it out because it was in the training set or prompt, yes. If it spits it out and it was in the training set (or prompt), it is at least difficult to make the case that it is not infringement (if it was the prompt, then you basically have the same problem as any other case where you are proving that production of something that matches something someone else has a copyright on was independent, since copyright only protects against copying, not coincidence.) If it spits it out but it was not in the training set or prompt then clearly it was not infringement.
> But if it ingests copyrighted material and then spits out entirely new content influenced by the originals… how is that an issue?
Because a non-perfect mechanical copy is still a copy violating copyright if there is no license, and a derivative work produced by a human using AI as a tool is still an infringing derivative work if there is no license. While copyright protects against perfect copies, it protects against imperfect copies and derivative works as well. (Except where exceptions like Fair Use apply.)
Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books
#68Earlier quoted context omitted.
I don't see much of a difference, seems like google profited off of copyrighted material that they didn't own. I dont understand the purpose of your feelings comment. I have no affiliations or preferences for any dogs in this fight. I don't necessarily think they are identical scenarios but if I were OpenAI's lawyer that's probably what I'd try and point at. I didnt say it would spit the book back out, at any point.…
Google isn't selling any books, nor is it distributing the book. OpenAI isn't pirating any books content nor is it distributing or reproducing it. Nor are they even "stealing the book". Both are transformative content.
But I understand what you're saying and it's certainly possible you're correct, I just don't think it's obvious.
Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books
#69Earlier quoted context omitted.
> Input vs output. This is what I've been wondering. Does Fair Use apply here at all? Sure, the models were trained on copyrighted material. But wouldn't the generative part of the AI count as transformative?
I was trained on copyrighted books. I only get in trouble when I spout out paragraphs from them from memory and pass them off as my own. I know not to do that, though. Seems only fair that the same should apply to GPT.
Honestly, this is starting to feel like copyright holders using the Big, New, Scary AI as a strawman to attack Fair Use.
Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books
#70Earlier quoted context omitted.
Because AI isn't human? It's a machine/algorithm?
Then why are these exceptions unique to human?
In the "monkey selfie" case, a photographer named David Slater set up a camera in the Indonesian jungle, and a macaque monkey took a photograph of itself with it. When the photo was uploaded and shared, various parties began to argue over who held the copyright. Slater claimed it was his because it was his camera and he set up the situation. Others believed that if the monkey pressed the shutter, then the monkey, or no one, held the copyright.
The U.S. Copyright Office clarified its stance on the matter in the Compendium of U.S. Copyright Office Practices, Third Edition. It stated:
"The U.S. Copyright Office will not register works produced by nature, animals, or plants. Likewise, the Office cannot register a work purportedly created by divine or supernatural beings, although the Office may register a work where the application or the deposit copy state that the work was inspired by a divine spirit."