Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

281–290 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#281

Earlier quoted context omitted.

I think the argument is somewhat more interesting if the book was pirated --both you as an individual and OpenAI as a company could be sued for that. But I really don't see how you could prove OpenAI did that, since ChatGPT could have learned from existing summaries on Wikipedia and Goodreads.

It seems pretty easy to prove that, since they admitted it in public. Read the article. This isn't about the question of LLMs being copyright infringement, this is about Meta and OpenAI admitting that they had pirated copies of those books.

> > But I really don't see how you could prove OpenAI did that

> It seems pretty easy to prove that, since they admitted it in public.

Can you highlight/link to where OpenAI have admitted this? As far as I'm aware, OpenAI are still secretive about their training datasets.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#282

I think this will be a bigger issue than some people think. Maybe there's a market for 'clean' training data that doesn't include potential copyright claims. Just public domain works. We'll know it's an AI because it talks like a late 18th century/early 19th century writer?

This isn't completely new, similar issues came up with search engines and this may be seen as 'transformative'. But there may be issues with models that happily reproduce copyrighted texts in their entirety along with other novel issues like models that hallucinate defamatory things or other such problems.

Still, I doubt this particular genie can be stuffed back into the bottle, so we'll probably see a lot of litigation and work on alignment, etc. along with new types of abuse.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#283

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

I’m imagining a world that looks just about the same as this one does. A larger book library doesn’t automatically make that medium more appealing to kids than what Mr Beast, Unspeakable, and the other crap kids love are doing.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#284

Earlier quoted context omitted.

That's not was he's saying at all. He's saying you can train an AI on copyrighted material just like people can learn from copyrighted material. If you acquire the material illegally that a separate issue that training AI doesn't give you any protection against.

Can I memorize copyrighted material and recite it on Youtube? What if I do so but imperfectly? Where do you draw the line? If it's infringement for a human to do that why is it not for a LLM?

That would be an unlicensed reproduction of the work. You’d need permission from the rights holder to create what is essentially an audiobook recording.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#285
> The lawsuit against OpenAI alleges that summaries of the plaintiffs’ work generated by ChatGPT indicate the bot was trained on their copyrighted content. “The summaries get some details wrong” but still show that ChatGPT “retains knowledge of particular works in the training dataset," the lawsuit says.

Setting aside the whole issue of whether LLM constitutes a derived work of whatever it's trained on, this sounds like a very weak argument to me. An LLM trained on numerous summaries of the works would also be capable of producing such summaries itself even if the works were never part of the training set. In general, having knowledge about something is not evidence of being trained on it.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#287

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Machine learning models have been trained with copyrighted data for a long time. Imagenet is full of copyrighted images, clearview literally just scanned the internet for faces, and I am sure there are other, older examples. I am unsure if this has been tested as fair use by a US court, but I am guessing it will be considered to be so if it is not already.

and not only copyrighted material, also illegal and disturbing content

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#288
post #83
post #38

>On information and belief, the reason ChatGPT can accurately summarize a certain copyrighted book is because that book was copied by OpenAI and ingested by the underlying OpenAI Language Model (either GPT-3.5 or GPT-4) as part of its training data. While it strikes me as perfectly plausible that the Books2 dataset contains Silverman's book, this quote from the complaint seems obviously false. First, even if the mode…

> this quote from the complaint seems obviously false I notice you go on to provide an argument only for why it might not be true. Also, seeing the other post on this, I asked chatgpt-4 for a summary of “ The Ruby of Kishmoor” as well, and it provided one to me, though I had to ask twice. I don’t know anything about that book, so I can’t tell if its summary is accurate, but so much for your test. It seems pretty naiv…

The test was whether producing equivalent accuracy and detail for summaries of all books in its training corpus was a feature of ChatGPT's ability to natively generate them from standalone source material or whether Silverman's detailed summary was likely just a "summary of summaries", not whether ChatGPT produced a result at all. From the comment you reference, it failed the test because the result was hallucinated.

You can pick something else that's in the training set that has SparkNotes and many popular reviews to compare. I routinely feed novel data sources into LLMs to test massive context and memory, and none produce anything similar in quality to what is being exhibited.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#289
post #173

Earlier quoted context omitted.

The point is that GP has no reason to believe Google, like Meta, also used copyrighted materials for training its AI. Why did Sarah Silverman sue OpenAI and Meta but not Google?

I didn't accuse Google of using copyrighted materials for training its AI. I accused Google of existing under a different set of laws than mere citizens. As an example, the mass usage of copyrighted materials to build Youtube.

[deleted]

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#290

Earlier quoted context omitted.

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

I’m imagining a world that looks just about the same as this one does. A larger book library doesn’t automatically make that medium more appealing to kids than what Mr Beast, Unspeakable, and the other crap kids love are doing.

...for the global middle class? Maybe.

For the world as a whole? Definite differences.

It just seems like a super jaded "kids these days" thing to hate on them for consuming easily accessed, free content- and acting like the global literacy and intellectual capital would remain unaffected.

Post reply on HN