Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

101–110 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#101

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification.

Aaron Swartz was a saint.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#102
post #61

Earlier quoted context omitted.

I'm allowed to make private copies of copywritten works. I'm not allowed to redistribute them. To what extent this is redistribution is not clear. Is there much of difference between this model and a machine, like a VCR, that recreates the original work when I press a button?

I buy a book and give it to my child, they read the book and later write and sell a story influenced by said book. should that be a copyright infringement? how about they become a therapist and sell access to knowledge from copyrighted books? should that be an infringement? what if they sell access to lectures they've given including facts from said book(s) to millions of people? it's understandable that people feel…

Making an analogy where you substitute a human being for the LLM is disingenuous to the extreme.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#103
post #92

Are we all reading the same complaint? They say: > in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private tracker.” Does that stack up? The Meta Paper -…

We don't seem to be reading the same thing, you're pulling Google out of thin air somewhere.

I'm literally quoting The Verge article and following the links they present...

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#104
post #7

I wonder if the makers of AI may have some index of learning material to perhaps prove they did not infringe? Or they just throw everything they get their hands on to the LLM?

It's the latter. The current approach is akin to throwing the entire internet at the foundational model, train it for an epoch or two, and that's that. Afterwards, the finetuning with curated training material, and techniques like RLHF (Reinforcement Learning from Human Feedback) takes place.

The sources of the training material are... questionable, to say the least. There's a reason the training dataset for GPT 3.5 and 4 remains undisclosed.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#105

Earlier quoted context omitted.

I can see a good argument in the complaint. The provenance of the training data leads back to it being acquired illegally. Illegally acquired materials were then used in a commercial venture. That the venture was an AI model is perhaps beside the point. You can’t use illegally acquired materials when doing business.

It seems like a weak argument, in that it is just as likely it saw any number of things about it, from book reviews to sales listings to interviews.

> it is just as likely it saw any number of things about it

Is this based on inside information, or just the law of averages? Doesn't the fact that they openly admitted to having been trained on pirated books affect your priors?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#106
post #102
post #61

Earlier quoted context omitted.

I buy a book and give it to my child, they read the book and later write and sell a story influenced by said book. should that be a copyright infringement? how about they become a therapist and sell access to knowledge from copyrighted books? should that be an infringement? what if they sell access to lectures they've given including facts from said book(s) to millions of people? it's understandable that people feel…

Making an analogy where you substitute a human being for the LLM is disingenuous to the extreme.

why do you think that?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#107

>> In the OpenAI suit, the trio offers exhibits showing that when prompted, ChatGPT will summarize their books, infringing on their copyrights. This doesn't seem like copyright infringement. I could read the book and offer a summary right? Someone on goodreads could as well. Why should an AI doing it be different? BTW I could also read someone's illicit copy and do the same, couldn't I? I think people are trying to c…

> I could read the book and offer a summary right? Someone on goodreads could as well. Why should an AI doing it be different?

We have to stop equating human beings to for-profit corporations running a machine at orders of magnitude the speed and scale. This is critical, otherwise we don’t have any arguments against – say a mass face recognition surveillance op because “humans can remember faces too”. Scale matters, just like it did before “AI” with things like indiscriminate surveillance “it’s just metadata” or “this location data set is just anonymized aggregates”.

> This doesn't seem like copyright infringement.

Now, I still think I agree with this. A book summary is nevertheless an extremely poor battle to pick, since frankly who the hell cares. It’s not like someone is gonna say “I’m not buying this book anymore because ChatGPT summarized it”.

Now, perhaps they just used the summary to prove that their book was part of the training set, and that they think it’s wrong to include their works without permission. That’s, imo, definitely not trivial to dismiss. Looks like unpaid supply chain to me.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#108

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

The lawsuit doesn't even mention Google.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#109
post #78
post #61

Earlier quoted context omitted.

I buy a book and give it to my child, they read the book and later write and sell a story influenced by said book. should that be a copyright infringement? how about they become a therapist and sell access to knowledge from copyrighted books? should that be an infringement? what if they sell access to lectures they've given including facts from said book(s) to millions of people? it's understandable that people feel…

If you found a way to have a million children who could grow up in one day your analogy would be more apt. In that case you and your children would rightly be considered a threat.

did you read to the third analogy?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#110

Sarah's pov raises some questions for me regarding my own "training", there is a noteworthy part of who I am built upon the consumed music, books, movies, video games and tv shows that myself or people around me have pirated and shared with me. This part of me helped me in life appreciably, I could also say I profited because of it, helping me along my life in being likable, funny, relatable, with broad outlooks etc.…

if you are a robot that ripped off literally all the data in the world and now resells it in a repackaged form for its own profit, then yes, you can be sued. Talking about whistling a song is pretty absurd in this context.
Post reply on HN