Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

41–50 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#41

Earlier quoted context omitted.

I'm allowed to make private copies of copywritten works. I'm not allowed to redistribute them. To what extent this is redistribution is not clear. Is there much of difference between this model and a machine, like a VCR, that recreates the original work when I press a button?

It's legal to make a copy of something you own, however it's not legal to make a copy of something illicitly acquired, whether or not there's distribution involved.

It gets kinda dicey in jurisdictions that have media levy taxes.

The uploader might be breaking the law but not the downloader that stores the copyrighted material on levy-paid media.

I’ve heard inklings of this argument in Canada but can’t figure out what the current state of the art is: https://en.m.wikipedia.org/wiki/File_sharing_in_Canada

Then is a corporation’s internal use for the purpose of analysis considered “private use”? If there’s no redistribution/broadcasting, is it still non-commercial?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#42

Earlier quoted context omitted.

I'm allowed to make private copies of copywritten works. I'm not allowed to redistribute them. To what extent this is redistribution is not clear. Is there much of difference between this model and a machine, like a VCR, that recreates the original work when I press a button?

It's legal to make a copy of something you own, however it's not legal to make a copy of something illicitly acquired, whether or not there's distribution involved.

Companies ofc now circumvent that by selling you licenses, and not something you own.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#44

A more convincing exhibit would have been convincing ChatGPT to output some of the text verbatim, instead of a summary. Here's what I got when I tried: I'm sorry for the inconvenience, but as of my knowledge cutoff in September 2021, I don't have access to specific external databases, books, or the ability to pull in new information after that date. This means that I can't provide a verbatim quote from Sarah Silverma…

I tried making chatgpt output first paragpraph of lord of the rings before, it goes silent after first few words. Looks like the devs are filtering it out

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#45
post #38

>On information and belief, the reason ChatGPT can accurately summarize a certain copyrighted book is because that book was copied by OpenAI and ingested by the underlying OpenAI Language Model (either GPT-3.5 or GPT-4) as part of its training data. While it strikes me as perfectly plausible that the Books2 dataset contains Silverman's book, this quote from the complaint seems obviously false. First, even if the mode…

Accessibility? I've heard of Silverman but never Ruby of Kishmoor

More people discuss it, more people summarize on their personal or other sites, etc

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#46

I mean I’m no lawyer but this doesn’t strike me as a great example for infringement? Detailed summaries of books sounds like textbook transformative use. Especially in Silverman’s case, reducing her book to “facts” while eliminating artistic elements of her prose make it that much less of a direct substitute for the original work.

I can see a good argument in the complaint. The provenance of the training data leads back to it being acquired illegally. Illegally acquired materials were then used in a commercial venture. That the venture was an AI model is perhaps beside the point. You can’t use illegally acquired materials when doing business.

It seems like a weak argument, in that it is just as likely it saw any number of things about it, from book reviews to sales listings to interviews.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#47

This is actually quite interesting, as it's drawing a distinction between training material that can be accessed by anybody with a web browser (like anybody's blog), vs. training material that was "illegally-acquired... available in bulk via torrent systems." I don't think there's any reason why this would be a relevant legal distinction in terms of distributing an LLM -- blog authors weren't giving consent either. H…

> Or do the courts not really care at all how something is made?

One of the fair use factors, which until fairly recently was consistently held out as the most important fair use factor, is the effect on the commercial market for the original work. Accordingly, a court is more likely to find that something is fair use if there is effectively no commercial market for the original work, though the fact that something isn't actively being sold isn't dispositive (open source licenses have survived in appellate courts despite being free as in beer).

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#48
post #38

>On information and belief, the reason ChatGPT can accurately summarize a certain copyrighted book is because that book was copied by OpenAI and ingested by the underlying OpenAI Language Model (either GPT-3.5 or GPT-4) as part of its training data. While it strikes me as perfectly plausible that the Books2 dataset contains Silverman's book, this quote from the complaint seems obviously false. First, even if the mode…

Accessibility? I've heard of Silverman but never Ruby of Kishmoor More people discuss it, more people summarize on their personal or other sites, etc

Right that is the point of the parent comment - it’s not the book, it’s the amalgamation of all the discussions and content about the book. This case is probably dead in the water.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#49
post #31

Earlier quoted context omitted.

This would be like you intensely studying the copy written work and then writing things based on the knowledge you obtained from that. Except, we don't know if their is an exception for things learned by people vs. things learned by machines, or if the machines are not really learning but copying instead (or if learning is intrinsically a form of copying?).

There's a sci-fi plot there: those with money can afford to pay the copyright cost for material they've learned and anything they produce results in royalties to the creators of everything they've learned. Those without means are cast out, perhaps some generating original thoughts in a way that breaks the system. I think I'm going to have to re-read the Unincorporated Man.

We aren't so far away from that, and it is becoming cheap enough that you won't even need money anymore. Rather than bootleg movies, people will just ask computers to derive a new movie from multiple existing movies, and then...it is an original work?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#50

This is actually quite interesting, as it's drawing a distinction between training material that can be accessed by anybody with a web browser (like anybody's blog), vs. training material that was "illegally-acquired... available in bulk via torrent systems." I don't think there's any reason why this would be a relevant legal distinction in terms of distributing an LLM -- blog authors weren't giving consent either. H…

Eventually, I imagine a new licensing concept will emerge, similar to the idea of music synchronization rights -- maybe call it "training rights." It won't matter whether the text was purchased or pirated -- just like it doesn't matter now if an audio track was purchased or pirated, when it's mixed into in a movie soundtrack.

Talent agencies will negotiate training rights fees in bulk for popular content creators, who will get a small trickle of income from LLM providers, paid by a fee line-itemed into the API cost. Indie creators' training rights will be violated willy-nilly, as they are now. Large for-profit LLMs suspected or proven as training rights violators will be shamed and/or sued. Indie LLMs will go under the radar.

Post reply on HN