Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

221–230 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#221

Earlier quoted context omitted.

LLMs are built upon neural networks which are modelled upon how brains work can you explain to me specifically how they're different? can you explain to me how they're different to the degree that making an analogy between the two is "disingenuous to the extreme"?

> LLMs are built upon neural networks which are modelled upon how brains work You are confused. Neural networks are inspired by how brains work, but they do not actually simulate brains. Airplanes are also inspired by how birds work, but (presumably) you don't think that bird laws should apply to airplanes. > can you explain to me how they're different to the degree that making an analogy between the two is "disingen…

can you point to the part of my comment where I suggested anything about how laws should work?

what do you think an analogy is? do you think an analogy is something where the thing itself and its metaphor are "just like" each other, or do they share attributes for illustrative purposes?

would you agree that LLMs and humans share the attributes of information retention and contextual reproduction?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#222
post #164

Earlier quoted context omitted.

Excellent, so you're saying I'll be able to download any copyrighted work from any pirate site and be free of all consequence if I just claim that I'm training an AI?

You can currently download any copyright work from any pirate site and be free of all consequences, and this has always been so. You just can't upload, since that counts as distribution, triggering civil and criminal penalties written in an age before the Internet when only shady commercial operators would distribute unlicensed copyrighted works.

Possession of copyrighted material without permission of the creator is illegal, but yes -- rights holders don't really go after infringers except maybe via ISP three strikes crap.

They're very much incentivized to change their behavior for AI scraping, though.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#223

Earlier quoted context omitted.

That's not was he's saying at all. He's saying you can train an AI on copyrighted material just like people can learn from copyrighted material. If you acquire the material illegally that a separate issue that training AI doesn't give you any protection against.

>you can train an AI on copyrighted material just like people can learn from copyrighted material. No amount of whining and hand wringing from engineers will ever make this true. This is for the courts to decide. A reasonable interpretation, in my eyes, is that the training process is a black box which takes in copyrighted works and produces a training model. The training model is a derivative work of the inputs. It…

[deleted]

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#224
post #176

Earlier quoted context omitted.

But acquiring the material illegally is the thrust of the issue, and the point of the thread here: the notion that large companies can get away with piracy if they just execute it on a large enough scale. Copyright infringment for thee but not for me.

There's no "for thee but not for me" issue here: nobody has ever been sued or prosecuted simply for downloading, acquiring, or possessing illegally acquired copyrighted works. People are sued and prosecuted for unlicensed distribution . Making and having your own copies, and doing what you want with them, has always been fine. At worst it's a grey area, but in many cases it's been protected as fair use.

I am not sure where you're getting your information from, but you're not a lawyer, so you shouldn't be so confident in telling people what things are fine legally.

Whether people have been sued for downloading works they don't have the right to copy onto their machines is irrelevant to whether it is actually illegal. And it certainly has nothing to do with fair use, which is about copyrighted works that you actually do have some right to.

IANAL so I'm not going to tell anyone what does and what does not constitute fair use in what jurisdictions.

BTW people are being sued for distribution because they make great examples because their offenses and thus the damages are much greater.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#225

Earlier quoted context omitted.

That's not was he's saying at all. He's saying you can train an AI on copyrighted material just like people can learn from copyrighted material. If you acquire the material illegally that a separate issue that training AI doesn't give you any protection against.

>you can train an AI on copyrighted material just like people can learn from copyrighted material. No amount of whining and hand wringing from engineers will ever make this true. This is for the courts to decide. A reasonable interpretation, in my eyes, is that the training process is a black box which takes in copyrighted works and produces a training model. The training model is a derivative work of the inputs. It…

I think a reasonable interpretation is also that what you are saying is correct, that doing all that does indeed infringe others' copyright, but that a fair use defense is valid.

I won't be particularly thrilled if that turns out to be the case, but I wouldn't be surprised if it does.

But as you say, we won't know until it's tested in court. And even then, often court cases around a complex topic like this will end up with a ruling that only clarifies a narrow aspect of it. So it might take many related court cases before we have a pretty good understanding of where the law stands. And then, of course, the law could change.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#226

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

> If you downloaded a book from that website, you would be sued and found guilty of infringement. How often does this actually happen? You might get handed an infringement notice, and your ISP might terminate your service if you're really egregious about it, but I haven't ever heard of someone actually being sued for downloading something.

[deleted]

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#227
post #92

Are we all reading the same complaint? They say: > in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private tracker.” Does that stack up? The Meta Paper -…

> But if Silverman's book is in there

It is:

    $ grep -i "Sarah Silverman" books3.list.txt
         325196 books3/the-eye.eu/public/Books/Bibliotik/T/The Bedwetter - Sarah Silverman.epub.txt
Anyone that just wants to see the list of files (itself a big file): https://gist.githubusercontent.com/Q726kbXuN/e4e9919a2f5d81f...

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#228

Could one argue that training AI systems constitutes an educational purpose, invoking the copyright exemption?

If you are making a commercial product I think not. Or do you mean educational as in educating the AI itself?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#229

Earlier quoted context omitted.

Piracy of content is against the law. All other analogies such as looking at paintings are not at issue here. The content was pirated and there are laws against that, whether we agree with it or not. So, if the plaintiff can prove the content was pirated, then the use of that content downstream is tainted.

>then the use of that content downstream is tainted What does that mean exactly? That's why I used the "looking at a stolen painting" example. Sure, pirating materials is illegal. But I don't think that's the big implication that people are getting at here. Is it legal to sell original works derived from perceiving stolen materials? Seems to me that it is.

In this case the correct analogy would be you brought a stolen painting into your house, looked at it for a while, and then produced your derivative work.

Surely you see the issue here? Receiving stolen property?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#230

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

> If you downloaded a book from that website, you would be sued and found guilty of infringement.

Suppose I buy a copy of a book, but then I spill my drink in it and it's ruined. If I go to the library, borrow the same book and make a photocopy of it to replace the damaged one I own, that might be fair use. Let's say for sake of argument that it is.

If instead I got the replacement copy from a piracy website, are you sure that's different?

Post reply on HN