Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

91–100 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#91
post #66

Earlier quoted context omitted.

Right, but that's useless without knowing how much (if any!) of it is actually correct. Is this completely hallucinated garbage?

How is it different from asking to me to summarize anything? I could have bought the book, or read the Wikipedia page, or listened people talking about it, or downloaded the torrent. In all those cases my summary could be right or could be wrong. If the rights holders know that I dowloaded the torrent they could sue me. In the other cases they can't. What if it turns out that OpenAI bought a copy of every book ingest…

> What if it turns out that OpenAI bought a copy of every book ingested be ChatGPT?

well, let's see the receipts then, they will surely have no problem winning in that case.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#92
Are we all reading the same complaint?

They say:

> in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private tracker.”

Does that stack up?

The Meta Paper - https://arxiv.org/pdf/2302.13971.pdf - says:

> We include two book corpora in our training dataset: the Gutenberg Project, which contains books that are in the public domain, and the Books3 section of ThePile (Gao et al., 2020)

The Pile Paper - https://arxiv.org/abs/2101.00027 - says it was trained (in part) on "Books3" which it describes as:

> Books3 is a dataset of books derived from a copy of the contents of the Bibliotik private tracker made available by Shawn Presser (Presser, 2020).

Shawn Presser's link is at https://twitter.com/theshawwn/status/1320282149329784833 and he describes Book3 as

> Presenting "books3", aka "all of bibliotik" - 196,640 books - in plain .txt

I don't have the time and space to download the 37GB file. But if Silverman's book is in there... isn't this a slam dunk case?

Meta's LLaMA is - as they seem to admit - trained on pirated books.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#93

Sarah's pov raises some questions for me regarding my own "training", there is a noteworthy part of who I am built upon the consumed music, books, movies, video games and tv shows that myself or people around me have pirated and shared with me. This part of me helped me in life appreciably, I could also say I profited because of it, helping me along my life in being likable, funny, relatable, with broad outlooks etc.…

> Is my brain just by the act of existing continually infringes on copyright? Can I be sued because I made a reference to a movie I pirated or because I whistle a song I never bought?

I think it would fall under fair use. But you can imagine what the world can become with microphones and cameras everywhere, which can already run music and speech recognition by themselves, in seconds. What a time to be alive!

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#94
post #50

This is actually quite interesting, as it's drawing a distinction between training material that can be accessed by anybody with a web browser (like anybody's blog), vs. training material that was "illegally-acquired... available in bulk via torrent systems." I don't think there's any reason why this would be a relevant legal distinction in terms of distributing an LLM -- blog authors weren't giving consent either. H…

Eventually, I imagine a new licensing concept will emerge, similar to the idea of music synchronization rights -- maybe call it "training rights." It won't matter whether the text was purchased or pirated -- just like it doesn't matter now if an audio track was purchased or pirated, when it's mixed into in a movie soundtrack. Talent agencies will negotiate training rights fees in bulk for popular content creators, wh…

>> Talent agencies will negotiate training rights fees in bulk for popular content creators

AFAICT there is no legal recognition of "training rights" or anything similar. First sale right is a thing, but even textbooks don't get extra rights for their training or educational value.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#95
post #66

Earlier quoted context omitted.

Right, but that's useless without knowing how much (if any!) of it is actually correct. Is this completely hallucinated garbage?

How is it different from asking to me to summarize anything? I could have bought the book, or read the Wikipedia page, or listened people talking about it, or downloaded the torrent. In all those cases my summary could be right or could be wrong. If the rights holders know that I dowloaded the torrent they could sue me. In the other cases they can't. What if it turns out that OpenAI bought a copy of every book ingest…

> What if it turns out that OpenAI bought a copy of every book ingested be ChatGPT?

Wouldn't that be trivial to prove if they had?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#97
post #62

Earlier quoted context omitted.

This refutes the previous post’s claim that chatgpt-4 refuses to even try to provide a summary.

Not necessarily, because the models have an element of randomness. Also, I was under the impression that ChatGPT has more "safeguards" (manifesting as a refusal to answer questions) than the raw API.

I don’t doubt the poster was telling the truth when they said they asked for a summary of the book and didn’t get one.

It refutes the idea that chatgpt’s inability to provide a summary means it didn’t scan the original text: since it can provide a summary, the argument is entirely spurious.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#98

Earlier quoted context omitted.

Accessibility? I've heard of Silverman but never Ruby of Kishmoor More people discuss it, more people summarize on their personal or other sites, etc

Right that is the point of the parent comment - it’s not the book, it’s the amalgamation of all the discussions and content about the book. This case is probably dead in the water.

I'm not entirely up to speed on US law, but wouldn't OpenAI have to provide the court some kind of proof that they didn't use it in the training data during discovery?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#99
post #92

Are we all reading the same complaint? They say: > in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private tracker.” Does that stack up? The Meta Paper -…

We don't seem to be reading the same thing, you're pulling Google out of thin air somewhere.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#100
post #66

Earlier quoted context omitted.

How is it different from asking to me to summarize anything? I could have bought the book, or read the Wikipedia page, or listened people talking about it, or downloaded the torrent. In all those cases my summary could be right or could be wrong. If the rights holders know that I dowloaded the torrent they could sue me. In the other cases they can't. What if it turns out that OpenAI bought a copy of every book ingest…

> In all those cases my summary could be right or could be wrong. Well that's incredibly nihilistic. Whether the summary is correct or not matters a great deal! And if someone I knew said they read a book, even a very obscure one, and then summarized it to me, I'd have great confidence that they would get such simple facts as "who are the characters" and "what are the major plot points" correct. But ChatGPT? Who the…

if someone I knew said they read a book, even a very obscure one, and then summarized it to me, I'd have great confidence that they would get such simple facts as "who are the characters" and "what are the major plot points" correct.

People, especially people you know, have reputations, based on history and experience that others have dealing with them. People can be known as liars, and anything they say is colored by such a reputation. Humans have language idioms for communicating about and dealing with such people too, phrases like "take anything that person says with a grain of salt". Look at how George Santos' history of lying about his own experience is being dealt with.

ChatGPT can be (is?) the same, and it has a bad reputation for truth telling. And LLMs' reputation is not necessarily getting better in this regard.

The problem is that many people attribute output that came from a machine to be of higher quality (on whatever axis) than output that came from a human, even a human they personally know and have experience dealing with. This is the same kind of prejudice as any other, or perhaps a more insidious prejudice.

Post reply on HN