Earlier quoted context omitted.
I'm allowed to make private copies of copywritten works. I'm not allowed to redistribute them. To what extent this is redistribution is not clear. Is there much of difference between this model and a machine, like a VCR, that recreates the original work when I press a button?
This would be like you intensely studying the copy written work and then writing things based on the knowledge you obtained from that. Except, we don't know if their is an exception for things learned by people vs. things learned by machines, or if the machines are not really learning but copying instead (or if learning is intrinsically a form of copying?).
Sarah Silverman is suing OpenAI and Meta for copyright infringement
31–40 of 599 posts
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#32I mean I’m no lawyer but this doesn’t strike me as a great example for infringement? Detailed summaries of books sounds like textbook transformative use. Especially in Silverman’s case, reducing her book to “facts” while eliminating artistic elements of her prose make it that much less of a direct substitute for the original work.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#33She'd have to sue every student that writes an essay on a book they'd read
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#34Earlier quoted context omitted.
This would be like you intensely studying the copy written work and then writing things based on the knowledge you obtained from that. Except, we don't know if their is an exception for things learned by people vs. things learned by machines, or if the machines are not really learning but copying instead (or if learning is intrinsically a form of copying?).
In the case of unreleased work, you writing about your knowledge of it is just proof that you obtained the work, which is proof that you committed a tort/trespass. Just like if you published a newspaper article with information you could only have acquired by hacking someone's phone. I'm not sure what a court would find against you, but it seems clear that there would be some way to couch that as a legal grievance.
A search engine that indexes the internet might be equally liable at that point, although the DCMA gives them an out if they have a mechanism to remove pirated entries from their index on request. Could LLMs have the same out?
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#35Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#36This is actually quite interesting, as it's drawing a distinction between training material that can be accessed by anybody with a web browser (like anybody's blog), vs. training material that was "illegally-acquired... available in bulk via torrent systems." I don't think there's any reason why this would be a relevant legal distinction in terms of distributing an LLM -- blog authors weren't giving consent either. H…
My understanding (disclaimer: IANAL) is that in order to claim fair use, you have to be legally in possession of the work. If the work is only legally available for sale, then you must have legally purchased a copy, or been given it by someone who did so (for example, if you received it as a gift).
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#37This is actually quite interesting, as it's drawing a distinction between training material that can be accessed by anybody with a web browser (like anybody's blog), vs. training material that was "illegally-acquired... available in bulk via torrent systems." I don't think there's any reason why this would be a relevant legal distinction in terms of distributing an LLM -- blog authors weren't giving consent either. H…
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#38While it strikes me as perfectly plausible that the Books2 dataset contains Silverman's book, this quote from the complaint seems obviously false.
First, even if the model never saw a single word of the book's text during training, it could still learn to summarize it from reading other summaries which are publicly available. Such as the book's Wikipedia page.
Second, it's not even clear to me that a model which only saw the text of a book, but not any descriptions or summaries of it, during training would even be particular good at producing a summary.
We can test this by asking for a summary of a book which is available through Project Gutenberg (which the complaint asserts is Books1 and therefore part of ChatGPT's training data) but for which there is little discussion online. If the source of the ability to summarize is having the book itself during training, the model should be equally able to summarize the rare book as it is Silverman's book.
I chose "The Ruby of Kishmoor" at random. It was added to PG in 2003. ChatGPT with GPT-3.5 hallucinates a summary that doesn't even identify the correct main characters. The GPT-4 model refuses to even try, saying it doesn't know anything about the story and it isn't part of its training data.
If ChatGPT's ability to summarize Silverman's book comes from the book itself being part of the training data, why can it not do the same for other books?
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#39Earlier quoted context omitted.
I'm allowed to make private copies of copywritten works. I'm not allowed to redistribute them. To what extent this is redistribution is not clear. Is there much of difference between this model and a machine, like a VCR, that recreates the original work when I press a button?
It's legal to make a copy of something you own, however it's not legal to make a copy of something illicitly acquired, whether or not there's distribution involved.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#40Isn’t it much more likely that there are a lot of book reviews and summaries in its training set from which it can synthesize its own?
Where are the summaries from? I would say it's much more likely that a shadow library was scraped but if course that is also seemingly impossible to prove. One may be able to somewhat test that by asking for a summary of a book/text only available on a shadow library.
As a step towards a PoC I looked at https://www.amazon.co.uk/Bedwetter-Stories-Courage-Redemptio... and found a reference to "Boys' Market Manchester" which seemed like a Googlewhack-ish (unlikely) triple of words. Then I asked ChatGPT about it:
Me: Has Sarah Silverman ever written about Boys' Market Manchester ChatGPT
ChatGPT: As of my knowledge cutoff in September 2021, I do not have any information indicating that Sarah Silverman has written specifically about Boys' Market Manchester. Sarah Silverman is an American comedian, actress, and writer known for her stand-up comedy and her work in film and television. While she has written books and has often shared personal anecdotes in her comedy, I couldn't find any specific references to Boys' Market Manchester in relation to her work. However, please note that my information might not be up to date, as Sarah Silverman's career and activities may have evolved since then.