>On information and belief, the reason ChatGPT can accurately summarize a certain copyrighted book is because that book was copied by OpenAI and ingested by the underlying OpenAI Language Model (either GPT-3.5 or GPT-4) as part of its training data. While it strikes me as perfectly plausible that the Books2 dataset contains Silverman's book, this quote from the complaint seems obviously false. First, even if the mode…
Sarah Silverman is suing OpenAI and Meta for copyright infringement
81–90 of 599 posts
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#82Sarah's pov raises some questions for me regarding my own "training", there is a noteworthy part of who I am built upon the consumed music, books, movies, video games and tv shows that myself or people around me have pirated and shared with me. This part of me helped me in life appreciably, I could also say I profited because of it, helping me along my life in being likable, funny, relatable, with broad outlooks etc.…
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#83>On information and belief, the reason ChatGPT can accurately summarize a certain copyrighted book is because that book was copied by OpenAI and ingested by the underlying OpenAI Language Model (either GPT-3.5 or GPT-4) as part of its training data. While it strikes me as perfectly plausible that the Books2 dataset contains Silverman's book, this quote from the complaint seems obviously false. First, even if the mode…
I notice you go on to provide an argument only for why it might not be true.
Also, seeing the other post on this, I asked chatgpt-4 for a summary of “ The Ruby of Kishmoor” as well, and it provided one to me, though I had to ask twice. I don’t know anything about that book, so I can’t tell if its summary is accurate, but so much for your test.
It seems pretty naive to me to just kind of assume chatgpt must be respecting copyright, and hasn’t scanned copyrighted material without obtaining authorization. Perhaps discovery will settle it, though. Logs of what they scanned should exist. (IMO, a better argument is that this is fair use.)
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#84>On information and belief, the reason ChatGPT can accurately summarize a certain copyrighted book is because that book was copied by OpenAI and ingested by the underlying OpenAI Language Model (either GPT-3.5 or GPT-4) as part of its training data. While it strikes me as perfectly plausible that the Books2 dataset contains Silverman's book, this quote from the complaint seems obviously false. First, even if the mode…
The playground summarizes it as this via GPT-4: Prompt: Please summarize the following book found on Project Gutenberg The Ruby of Kishmoor Response: "The Ruby of Kishmoor" is a short adventure story written by Howard Pyle. The narrative revolves around the life of Jonathan Rugg, a young man who is enticed by a mysterious stranger to come to the Caribbean to secure a valuable relic, the Ruby of Kishmoor. Once Jonatha…
[0] https://www.gutenberg.org/cache/epub/3687/pg3687-images.html
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#85Intellectual property in the way of progress, again.
Lawsuits like this are tests to evaluate the current state of affairs and to force legislation into dealing with the greater issue of AI in context of copyright, IP, and fair use. It would only be "in the way" if it would actually stop or hinder anything, which a lawsuit on its own isn't.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#86This is the makers of AI explicitly saying that they did use copyrighted works from a book piracy website. If you downloaded a book from that website, you would be sued and found guilty of infringement. If you downloaded all of them, you would be liable for many billions of dollars in damages.
But companies like Google and Facebook get to play by different rules. Kill one person and you're a murderer, kill a million and to ask you about it is a "gotcha question" that you can react to with outrage.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#87>> In the OpenAI suit, the trio offers exhibits showing that when prompted, ChatGPT will summarize their books, infringing on their copyrights. This doesn't seem like copyright infringement. I could read the book and offer a summary right? Someone on goodreads could as well. Why should an AI doing it be different? BTW I could also read someone's illicit copy and do the same, couldn't I? I think people are trying to c…
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#88>On information and belief, the reason ChatGPT can accurately summarize a certain copyrighted book is because that book was copied by OpenAI and ingested by the underlying OpenAI Language Model (either GPT-3.5 or GPT-4) as part of its training data. While it strikes me as perfectly plausible that the Books2 dataset contains Silverman's book, this quote from the complaint seems obviously false. First, even if the mode…
The playground summarizes it as this via GPT-4: Prompt: Please summarize the following book found on Project Gutenberg The Ruby of Kishmoor Response: "The Ruby of Kishmoor" is a short adventure story written by Howard Pyle. The narrative revolves around the life of Jonathan Rugg, a young man who is enticed by a mysterious stranger to come to the Caribbean to secure a valuable relic, the Ruby of Kishmoor. Once Jonatha…
The plot of the story is that Jonathan Rugg is a Quaker who works as a clerk in Philadelphia. His boss sends him on a trip to Jamaica (credit for mentioning the Caribbean!). After arriving, he meets a woman who asks him to guard for her an ivory ball, and says that there are three men after her who want to steal it. By coincidence, he runs into the first man, they talk, he shows him the ball, and the man pulls a knife. In the struggle, the man is accidentally stabbed. Another man arrives, and sees the scene. Jonathan tries to explain, and shows him the orb. The man pulls a gun, and in the struggle is accidentally shot. A third man arrives, same story, they go down to the dock to dispose of the bodies and the man tries to steal the orb. In the struggle he is killed by when a plank of the dock collapses. Jonathan returns to the woman and says he has to return the orb to her because it's brought too much trouble. She says the men who died were the three after her, and reveals that the orb is actually a container, holding the ruby. She offers to give him the ruby and to marry him. He refuses, saying that he is already engaged back in Philadelphia, and doesn't want anything more to do with the ruby. He returns to Philadelphia and gets married, swearing off any more adventures.
https://en.wikisource.org/wiki/Howard_Pyle%27s_Book_of_Pirat...
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#89Earlier quoted context omitted.
Right, but that's useless without knowing how much (if any!) of it is actually correct. Is this completely hallucinated garbage?
How is it different from asking to me to summarize anything? I could have bought the book, or read the Wikipedia page, or listened people talking about it, or downloaded the torrent. In all those cases my summary could be right or could be wrong. If the rights holders know that I dowloaded the torrent they could sue me. In the other cases they can't. What if it turns out that OpenAI bought a copy of every book ingest…
Producing the summary is absolutely not an infringing act. Downloading the torrent might be.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#90This is actually quite interesting, as it's drawing a distinction between training material that can be accessed by anybody with a web browser (like anybody's blog), vs. training material that was "illegally-acquired... available in bulk via torrent systems." I don't think there's any reason why this would be a relevant legal distinction in terms of distributing an LLM -- blog authors weren't giving consent either. H…
> Is there any legal basis for saying fair use permits distributing an LLM trained on copyrighted material, but you have to purchase all the content first to do so legally if it's only available for sale? My understanding (disclaimer: IANAL) is that in order to claim fair use, you have to be legally in possession of the work. If the work is only legally available for sale, then you must have legally purchased a copy,…
I am also NAL, but can I imagine it goes further than that. Just purchasing a copy doesn't let you create and sell (directly as content or indirectly via a service like a chatbot) derivative works that are substantially similar in style and voice to the original work.
For example, an LLM 's response to the request:
"Write a short story about a comical trip to the nail salon in the style of Sarah Silverman"
... IMO doesn't constitute fair use, because the intellectual property of the artist is their style even more than the content they produce. Their style, built from their lived human experience, is what generates their copyrighted content. Even more than the content, the artist's style should be protected. The fact that a technology exists that can convincingly mimic their style doesn't change that.
One might then ask, well what about artists mimicking each others work? Well, any artist with a shred of integrity will credit their major influences.
We should hold machines (and their creators) to an even tougher standard than we hold people when it comes to mimicry. A real person can be inspired and moved by another person's artistic work such that they mimic it. Inspiration means nothing to a machine.