Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

31–40 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#31

Earlier quoted context omitted.

I'm allowed to make private copies of copywritten works. I'm not allowed to redistribute them. To what extent this is redistribution is not clear. Is there much of difference between this model and a machine, like a VCR, that recreates the original work when I press a button?

This would be like you intensely studying the copy written work and then writing things based on the knowledge you obtained from that. Except, we don't know if their is an exception for things learned by people vs. things learned by machines, or if the machines are not really learning but copying instead (or if learning is intrinsically a form of copying?).

There's a sci-fi plot there: those with money can afford to pay the copyright cost for material they've learned and anything they produce results in royalties to the creators of everything they've learned. Those without means are cast out, perhaps some generating original thoughts in a way that breaks the system. I think I'm going to have to re-read the Unincorporated Man.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#32

I mean I’m no lawyer but this doesn’t strike me as a great example for infringement? Detailed summaries of books sounds like textbook transformative use. Especially in Silverman’s case, reducing her book to “facts” while eliminating artistic elements of her prose make it that much less of a direct substitute for the original work.

I can see a good argument in the complaint. The provenance of the training data leads back to it being acquired illegally. Illegally acquired materials were then used in a commercial venture. That the venture was an AI model is perhaps beside the point. You can’t use illegally acquired materials when doing business.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#34

Earlier quoted context omitted.

This would be like you intensely studying the copy written work and then writing things based on the knowledge you obtained from that. Except, we don't know if their is an exception for things learned by people vs. things learned by machines, or if the machines are not really learning but copying instead (or if learning is intrinsically a form of copying?).

In the case of unreleased work, you writing about your knowledge of it is just proof that you obtained the work, which is proof that you committed a tort/trespass. Just like if you published a newspaper article with information you could only have acquired by hacking someone's phone. I'm not sure what a court would find against you, but it seems clear that there would be some way to couch that as a legal grievance.

Hmm...if I get on the internet and download and read a paper from some website, am I liable if that paper was actually private if I had no clue it was obtained illegally? It seems to me that the distributor would be liable at that point, not the person who got it from the distributor (unless they knew they were stolen goods, then of course they are liable!).

A search engine that indexes the internet might be equally liable at that point, although the DCMA gives them an out if they have a mechanism to remove pirated entries from their index on request. Could LLMs have the same out?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#36

This is actually quite interesting, as it's drawing a distinction between training material that can be accessed by anybody with a web browser (like anybody's blog), vs. training material that was "illegally-acquired... available in bulk via torrent systems." I don't think there's any reason why this would be a relevant legal distinction in terms of distributing an LLM -- blog authors weren't giving consent either. H…

> Is there any legal basis for saying fair use permits distributing an LLM trained on copyrighted material, but you have to purchase all the content first to do so legally if it's only available for sale?

My understanding (disclaimer: IANAL) is that in order to claim fair use, you have to be legally in possession of the work. If the work is only legally available for sale, then you must have legally purchased a copy, or been given it by someone who did so (for example, if you received it as a gift).

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#37

This is actually quite interesting, as it's drawing a distinction between training material that can be accessed by anybody with a web browser (like anybody's blog), vs. training material that was "illegally-acquired... available in bulk via torrent systems." I don't think there's any reason why this would be a relevant legal distinction in terms of distributing an LLM -- blog authors weren't giving consent either. H…

[deleted]

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#38
>On information and belief, the reason ChatGPT can accurately summarize a certain copyrighted book is because that book was copied by OpenAI and ingested by the underlying OpenAI Language Model (either GPT-3.5 or GPT-4) as part of its training data.

While it strikes me as perfectly plausible that the Books2 dataset contains Silverman's book, this quote from the complaint seems obviously false.

First, even if the model never saw a single word of the book's text during training, it could still learn to summarize it from reading other summaries which are publicly available. Such as the book's Wikipedia page.

Second, it's not even clear to me that a model which only saw the text of a book, but not any descriptions or summaries of it, during training would even be particular good at producing a summary.

We can test this by asking for a summary of a book which is available through Project Gutenberg (which the complaint asserts is Books1 and therefore part of ChatGPT's training data) but for which there is little discussion online. If the source of the ability to summarize is having the book itself during training, the model should be equally able to summarize the rare book as it is Silverman's book.

I chose "The Ruby of Kishmoor" at random. It was added to PG in 2003. ChatGPT with GPT-3.5 hallucinates a summary that doesn't even identify the correct main characters. The GPT-4 model refuses to even try, saying it doesn't know anything about the story and it isn't part of its training data.

If ChatGPT's ability to summarize Silverman's book comes from the book itself being part of the training data, why can it not do the same for other books?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#39

Earlier quoted context omitted.

I'm allowed to make private copies of copywritten works. I'm not allowed to redistribute them. To what extent this is redistribution is not clear. Is there much of difference between this model and a machine, like a VCR, that recreates the original work when I press a button?

It's legal to make a copy of something you own, however it's not legal to make a copy of something illicitly acquired, whether or not there's distribution involved.

is there legal jargon for this distinction?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#40

Isn’t it much more likely that there are a lot of book reviews and summaries in its training set from which it can synthesize its own?

Where are the summaries from? I would say it's much more likely that a shadow library was scraped but if course that is also seemingly impossible to prove. One may be able to somewhat test that by asking for a summary of a book/text only available on a shadow library.

You could ingest all reviews that are extent in the online corpus, remove from the book all quotes found. Then ask the AI if a distinctive triples, say, of words appeared in their book, somehow, you'd probably need prompt engineering to get past "While I don't have access to the full text of the book [...]". A little maths and you might prove beyond reasonable doubt that the LLM was trained on the book.

As a step towards a PoC I looked at https://www.amazon.co.uk/Bedwetter-Stories-Courage-Redemptio... and found a reference to "Boys' Market Manchester" which seemed like a Googlewhack-ish (unlikely) triple of words. Then I asked ChatGPT about it:

Me: Has Sarah Silverman ever written about Boys' Market Manchester ChatGPT

ChatGPT: As of my knowledge cutoff in September 2021, I do not have any information indicating that Sarah Silverman has written specifically about Boys' Market Manchester. Sarah Silverman is an American comedian, actress, and writer known for her stand-up comedy and her work in film and television. While she has written books and has often shared personal anecdotes in her comedy, I couldn't find any specific references to Boys' Market Manchester in relation to her work. However, please note that my information might not be up to date, as Sarah Silverman's career and activities may have evolved since then.

Post reply on HN