Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

51–60 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#51
post #38

>On information and belief, the reason ChatGPT can accurately summarize a certain copyrighted book is because that book was copied by OpenAI and ingested by the underlying OpenAI Language Model (either GPT-3.5 or GPT-4) as part of its training data. While it strikes me as perfectly plausible that the Books2 dataset contains Silverman's book, this quote from the complaint seems obviously false. First, even if the mode…

The playground summarizes it as this via GPT-4:

Prompt:

Please summarize the following book found on Project Gutenberg

The Ruby of Kishmoor

Response:

"The Ruby of Kishmoor" is a short adventure story written by Howard Pyle. The narrative revolves around the life of Jonathan Rugg, a young man who is enticed by a mysterious stranger to come to the Caribbean to secure a valuable relic, the Ruby of Kishmoor.

Once Jonathan reaches the Caribbean, he learns of the ruby's grave curse. Despite this, his innate curiosity and the prospect of considerable wealth drive him to risk his life and go after the highly prized gem. After facing a slew of challenges and trials, he surprisingly stumbles upon the ruby through sheer luck rather than organized search efforts.

In the course of his journey, Jonathan delves into the depths of human greed and the lengths to which individuals go for materialistic gains. The story encapsulates an amalgamation of high-tide adventures, brushes with supernatural elements, noteworthy instances of bravery, and deep moral reflections, demanding readers to rethink about the real value of materialistic pursuits.

Jonathan finally manages to escape with the ruby but at great personal cost. The story ends on a sobering note, making readers wonder if the adventures and trials were worth the prize, thus engaging them in a discourse about the true value of life and material pursuits.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#52
Getty Images also filed an AI lawsuit, alleging that Stability AI ...

lol, bad karma? So it is okay for Getty to steal from others, but not ok for others to steal from them? I don't have a dog in this fight, but the goddamn the hypocrisy of these companies...

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#54

I mean I’m no lawyer but this doesn’t strike me as a great example for infringement? Detailed summaries of books sounds like textbook transformative use. Especially in Silverman’s case, reducing her book to “facts” while eliminating artistic elements of her prose make it that much less of a direct substitute for the original work.

Perhaps not, I thought one of the claims is interesting though, that they illegally acquired some of the dataset. What would be the damages from that, the retail price of the hardcopy?

The remedies under Title 17 are an injunction against further distribution, disgorgement or statutory damages, and potentially attorneys fees. The injunction part is why these cases usually settle if the defendant is actually in the wrong.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#55

Isn’t it much more likely that there are a lot of book reviews and summaries in its training set from which it can synthesize its own?

If book reviews and summaries were part of the training set, wouldn't that imply that OpenAI's LLM is more like a search engine in that it produces the input text based on a prompt?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#56
post #38

>On information and belief, the reason ChatGPT can accurately summarize a certain copyrighted book is because that book was copied by OpenAI and ingested by the underlying OpenAI Language Model (either GPT-3.5 or GPT-4) as part of its training data. While it strikes me as perfectly plausible that the Books2 dataset contains Silverman's book, this quote from the complaint seems obviously false. First, even if the mode…

The playground summarizes it as this via GPT-4: Prompt: Please summarize the following book found on Project Gutenberg The Ruby of Kishmoor Response: "The Ruby of Kishmoor" is a short adventure story written by Howard Pyle. The narrative revolves around the life of Jonathan Rugg, a young man who is enticed by a mysterious stranger to come to the Caribbean to secure a valuable relic, the Ruby of Kishmoor. Once Jonatha…

Right, but that's useless without knowing how much (if any!) of it is actually correct. Is this completely hallucinated garbage?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#57

Earlier quoted context omitted.

I can see a good argument in the complaint. The provenance of the training data leads back to it being acquired illegally. Illegally acquired materials were then used in a commercial venture. That the venture was an AI model is perhaps beside the point. You can’t use illegally acquired materials when doing business.

It seems like a weak argument, in that it is just as likely it saw any number of things about it, from book reviews to sales listings to interviews.

Unless OpenAI can prove that the outputs are derived from legally vs illegally-obtained outputs, not sure that’s going to matter. And as far as I understand about their models, that’s effectively impossible.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#58

Earlier quoted context omitted.

Accessibility? I've heard of Silverman but never Ruby of Kishmoor More people discuss it, more people summarize on their personal or other sites, etc

Right that is the point of the parent comment - it’s not the book, it’s the amalgamation of all the discussions and content about the book. This case is probably dead in the water.

> This case is probably dead in the water.

Is that a fact? I’m no lawyer, but if they can get it in front of a jurry is it impossible that they will find a human author more relatable and the technical counter arguments goblydook?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#59

Earlier quoted context omitted.

Perhaps not, I thought one of the claims is interesting though, that they illegally acquired some of the dataset. What would be the damages from that, the retail price of the hardcopy?

Wouldn't they first need to prove that OpenAI didn't ingest summaries of the book, and not the book itself?

[deleted]

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#60

This is actually quite interesting, as it's drawing a distinction between training material that can be accessed by anybody with a web browser (like anybody's blog), vs. training material that was "illegally-acquired... available in bulk via torrent systems." I don't think there's any reason why this would be a relevant legal distinction in terms of distributing an LLM -- blog authors weren't giving consent either. H…

>> blog authors weren't giving consent either.

That is a good point, since copyright is a default protection of works created by people.

Post reply on HN