Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

181–190 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#181

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

> If you downloaded a book from that website, you would be sued and found guilty of infringement. How often does this actually happen? You might get handed an infringement notice, and your ISP might terminate your service if you're really egregious about it, but I haven't ever heard of someone actually being sued for downloading something.

>How often does this actually happen?

Did you hear about Aaron Schwartz?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#182

Earlier quoted context omitted.

I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification. Aaron Swartz was a saint.

There's a difference between "information wants to be free" and "Facebook can produce works minimally derived from your greatest creative work at a scale you can't match". LLMs seem to aggregate that value to whoever builds the model, which they can then sell access to, or sell the output it produces. Five years from now, will OpenAI actually be open, or will it be a rent seeking org chasing the next quarterly gains?…

"Will OpenAI actually be open"

That ship sailed, friend.

OpenAI is no longer a charity in any meaningful sense of the word anymore, it's now an adversarial organization working against the public good with the sole aim of making a few rich men richer.

After privatization, they sent their PR people to lobby congress to make it impossible for anyone to compete with them (important note: not out of any interest in actually "protecting" the public from the very AI they're building), and perhaps worst of all, they're no longer being open with the scientific theories and data behind their new models.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#183
post #149

Earlier quoted context omitted.

The differences between a a human being and a computer are too numerous to list. I don’t even know why you need to ask the question. Let me ask another question to point out the absurdity of yours: Human beings have more in common with a bacterium than a software program. Can you tell me specifically how humans are not bacteria?

analogies are not descriptions of the things themselves, otherwise they would not be analogies, would they? now remember that this is an analogy . re-read my comments in this light and perhaps we can continue this conversation in a more grounded and reasonable manner however, I'll be frank: have you studied neural networks? if you haven't, it's very difficult to take you seriously on this

Now you’re just being condescending. I did read your comments and I know what an analogy is. Consider that there is a different perspective to yours that can validly view your analogy as absurd.

Yes I have studied neural nets and have a good understanding of their function. I am still not sure how, despite their development being inspired by animal brains, you can liken an LLM to an actual person. There are so many vast differences. Do you really want me to explain specially what they are? Surely, since both you and I are so familiar with the subject matter, that is unnecessary.

If we were taking about an AGI then this would be an entirely different conversation.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#184
post #139

Earlier quoted context omitted.

Because LLMs are not people. They are nothing like people; not in construction nor behaviour.

LLMs are built upon neural networks which are modelled upon how brains work can you explain to me specifically how they're different? can you explain to me how they're different to the degree that making an analogy between the two is "disingenuous to the extreme"?

[deleted]

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#185

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

I've pirated many books, never sued.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#186

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

I've pirated many books, never sued.

You mean, never caught.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#187
post #176

Earlier quoted context omitted.

That's not was he's saying at all. He's saying you can train an AI on copyrighted material just like people can learn from copyrighted material. If you acquire the material illegally that a separate issue that training AI doesn't give you any protection against.

But acquiring the material illegally is the thrust of the issue, and the point of the thread here: the notion that large companies can get away with piracy if they just execute it on a large enough scale. Copyright infringment for thee but not for me.

Large countries seem to be doing it as well - https://petapixel.com/2023/06/05/japan-declares-ai-training-...

And as long as OpenAI have an office in Japan they can absolutely legally train the models, no?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#188

Earlier quoted context omitted.

>You can’t use illegally acquired materials when doing business. This vague sentence conjures images of a company building products from stolen parts, but this situation seems different. IANAL, but if I looked at a stolen painting that nobody had ever seen, and sold handwritten descriptions of the painting to whoever wanted to buy one, I'm pretty sure what I've sold is not illegal.

Piracy of content is against the law. All other analogies such as looking at paintings are not at issue here. The content was pirated and there are laws against that, whether we agree with it or not. So, if the plaintiff can prove the content was pirated, then the use of that content downstream is tainted.

>then the use of that content downstream is tainted

What does that mean exactly? That's why I used the "looking at a stolen painting" example.

Sure, pirating materials is illegal. But I don't think that's the big implication that people are getting at here. Is it legal to sell original works derived from perceiving stolen materials? Seems to me that it is.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#189

If they ripped all of Bibliotik, the more interesting story to me is how they were able to get it all without hitting ratio requirements? Super fast internet that downloaded all they could before being ratio banned, overwhelmingly fast internet that was hopping on all the popular torrents to slowly build up ratio?

Can easily buy compromised accounts for most torrent sites on the dark web.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#190

Earlier quoted context omitted.

Right that is the point of the parent comment - it’s not the book, it’s the amalgamation of all the discussions and content about the book. This case is probably dead in the water.

I'm not entirely up to speed on US law, but wouldn't OpenAI have to provide the court some kind of proof that they didn't use it in the training data during discovery?

No. Burden is on the plaintiff (Silverman) to prove infringement.
Post reply on HN