Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

371–380 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#371
post #362

Earlier quoted context omitted.

I think the onus should be on Sarah Silverman to prove which inputs were hers and which outputs leveraged those inputs. I think she should pay all of the court costs if she fails to do so.

Well yeah - that's how the legal system works assuming it gets all the way to court. In reality, Meta and OpenAIs lawyers will do a risk evaluation against the strength of the claim and if there's any merit at all there will be a quiet settlement.

IANAL, but my understanding is that in the United States, a civil defendant who prevails is not usually entitled to have their legal fees paid by the plaintiff.

[0] https://porterlaw.com/obtaining-attorney-fees-in-litigation-...

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#372

Earlier quoted context omitted.

> I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification. I would be very happy if either a court or lawmakers decided that copyright itself was unconscionable. That isn't what's going to happen, though. And I think it's incredibly unacceptable if a court…

You're analyzing the factors with respect to the output of the model, not the weights: > The purpose and character is absolutely heavily commercial and makes a great deal of money for the companies building the AIs. That's assuming that it's Microsoft/OpenAI. Suppose a non-profit trains a model and releases the weights for free. > There's nothing about the works used for AI training that makes them any less entitled…

> You can't find a copy of any specific work anywhere in the weights.

You won't find a copy of any specific work anywhere in the compressed form of a file, either, but when you decompress it you find the complete work. And many large AIs can recite, verbatim or near-verbatim, many complete works. Yes, they might get a word wrong, but that doesn't nullify the point that they're trained on the entire work and to a first approximation they can emit the whole work.

> Typically the consumers of the weights are software developers or content creators, whereas the consumers of the original text or image are fans.

Many of the consumers of image models are in fact generating art that they previously would have commissioned from an artist. (Some of them are also generating art they never would have commissioned from an artist, so I'm not implying that this is a one-for-one revenue loss.) Consumers of code models are, in fact, potentially reducing the total demand for novice programmers.

The model derived from a pile of artistic works is, in fact, directly competing with those artistic works.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#373

Earlier quoted context omitted.

I think the onus should be on Sarah Silverman to prove which inputs were hers and which outputs leveraged those inputs. I think she should pay all of the court costs if she fails to do so.

how could someone "prove" which inputs and outputs of a large ML model leveraged any specific data?

That's her and her legal team's homework assignment.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#374
post #83

Earlier quoted context omitted.

> this quote from the complaint seems obviously false I notice you go on to provide an argument only for why it might not be true. Also, seeing the other post on this, I asked chatgpt-4 for a summary of “ The Ruby of Kishmoor” as well, and it provided one to me, though I had to ask twice. I don’t know anything about that book, so I can’t tell if its summary is accurate, but so much for your test. It seems pretty naiv…

> IMO, a better argument is that this is fair use There is no way in Hell that this is fair use! Fair use defenses rest on the fact that a limited excerpt was used for limited distribution, among other criteria. For example, if I'm a teacher and I make 30 copies of one page of a 300-page novel and I hand that out to my students, that's a brief excerpt for a fairly limited distribution. Now if I'm a social media influ…

You are assuming the LLM spits out exact copies of everything they’ve read.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#375

Earlier quoted context omitted.

Not really. The "special" quality attached to humans is only in creating copyright -- it has nothing to do with fair use arguments around derivative works. "Machines and algorithms are not legally recognized as being able to author original non-derivative works" Neither are monkeys. This doesn't mean a monkey's painting is any more or less derivative, or any more or less subject to a copyright claim. It only means th…

Monkeys aren't algorithms nor computers, so that doesn't seem very relevant. Let's look at a totally different analogy: compression algorithms. If I take a digital artist's work which they publish as a png or psd file, and I use some algorithm to convert it to a jpg file, well, I definitely transformed the work in terms of bytes. It's a smaller file, I threw out a lot of data, you can't get the original back. Yet, th…

But this is only relevant to copyright - and more over, when it comes to derivative works, only relevant to commercial impact.

An LLM outputting a summary of someone's work (1) doesn't create a new copyright work (so no profit can be derived from it's sale) but (2) would fail the test of whether it was competing with the original copyright work.

i.e. no one looking for a summary of a comedy skit is then going to consume that in preference to consuming the original skit. If you tried to argue that was the case, you'd then have to answer why a human review or wikipedia summary does not consitute an infringement.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#376
post #36

Earlier quoted context omitted.

> Is there any legal basis for saying fair use permits distributing an LLM trained on copyrighted material, but you have to purchase all the content first to do so legally if it's only available for sale? My understanding (disclaimer: IANAL) is that in order to claim fair use, you have to be legally in possession of the work. If the work is only legally available for sale, then you must have legally purchased a copy,…

> in order to claim fair use, you have to be legally in possession of the work. Which work? The original work, or the derivative work that you're using? Wikipedia uses non-free content all the time, and they're not purchasing albums to do it. Wikipedia reduces album covers, for example, to low resolution, so that they could not be reused to reproduce a real cover, for example. Sometimes Wikipedia uses screencaps of a…

> Which work? The original work

Yes. You create the derivative work, which automatically means you are legally in possession of it--even if it breaks the law in other respects.

> Wikipedia uses non-free content all the time

How are they obtaining it? From websites where albums are advertised? The images on those websites are available to the public for free, even if the albums or the album covers are only available for sale.

Also, Wikipedia articles are contributed to by individual people, who might well own copies of books that they quote from in the articles, for example, even if the corporation that owns Wikipedia does not. AFAIK, by Wikipedia's terms of use, if you post content you are implicitly asserting that you have a legal right to post it, so if there were a lawsuit they would probably punt to whoever posted the content.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#377

Earlier quoted context omitted.

> In all those cases my summary could be right or could be wrong. Well that's incredibly nihilistic. Whether the summary is correct or not matters a great deal! And if someone I knew said they read a book, even a very obscure one, and then summarized it to me, I'd have great confidence that they would get such simple facts as "who are the characters" and "what are the major plot points" correct. But ChatGPT? Who the…

if someone I knew said they read a book, even a very obscure one, and then summarized it to me, I'd have great confidence that they would get such simple facts as "who are the characters" and "what are the major plot points" correct. People, especially people you know, have reputations, based on history and experience that others have dealing with them. People can be known as liars, and anything they say is colored b…

Agree 100% with all of this. LLMs have a huge reputation problem; you simply cannot trust what they say because they've been proven time and time again to hallucinate fictional answers. Until that problem is solved I'm struggling to see how they're as useful as people are claiming they are.

You know what would be a fun test of integrity -- look up an obscure novel (potentially even the aforementioned one) that you know LLMs consistently hallucinate about because the details aren't in its training set, and then assign an essay about it as an academic assignment. It'll be pretty obvious who's read the book and who merely consulted an LLM because the latter will just be complete gibberish to anyone who actually knows what happens in that novel.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#378

Earlier quoted context omitted.

There's an interesting nuance here if you were to put a human in the place of the LLM. We have read thousands of works; does that mean anything we write is derivative?

Humans are special and can create new copyrights. The process of a human brain synthesizing stuff does act as a barrier to copyright infringement. Machines and algorithms are not legally recognized as being able to author original non-derivative works. > put a human in the place of the LLM But also, no, if you have a team of humans doing rote matrix multiplication instead of an LLM, that does not make it so the matri…

> Machines and algorithms are not legally recognized as being able to author original non-derivative works.

“original non-derivative” is noise: only humans can author works. This is especially equally true of derivative works, which must themselves being distinct works of authorship (a mechanical copy is not a derivative work, its a copy.)

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#379

Earlier quoted context omitted.

how could someone "prove" which inputs and outputs of a large ML model leveraged any specific data?

That's her and her legal team's homework assignment.

no, it isn't

it's an unsatisfiable requirement, and unnecessary to substantiate the legal claims

it's dumb to talk about

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#380

Earlier quoted context omitted.

Not really. The "special" quality attached to humans is only in creating copyright -- it has nothing to do with fair use arguments around derivative works. "Machines and algorithms are not legally recognized as being able to author original non-derivative works" Neither are monkeys. This doesn't mean a monkey's painting is any more or less derivative, or any more or less subject to a copyright claim. It only means th…

Monkeys aren't algorithms nor computers, so that doesn't seem very relevant. Let's look at a totally different analogy: compression algorithms. If I take a digital artist's work which they publish as a png or psd file, and I use some algorithm to convert it to a jpg file, well, I definitely transformed the work in terms of bytes. It's a smaller file, I threw out a lot of data, you can't get the original back. Yet, th…

> Is there a way that an LLM isn't, legally, a compression algorithm for a large set of copyrighted works?

An LLM is a lossy compression algorithm for a body of data (which may or may not consists of multiple “works” under copyright, and any works included in the data may or may not be protected by copyright), which body as a whole likely comprises a work (as compilation) which may or may not legally be derivative of some or all copyright protected works contained in the compilation, before considering Fair Use analysis.

It is not particularly a compression algorithm for the individual works if the body of data consists of individual works.

Post reply on HN