Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

251–260 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#251

Earlier quoted context omitted.

That's not was he's saying at all. He's saying you can train an AI on copyrighted material just like people can learn from copyrighted material. If you acquire the material illegally that a separate issue that training AI doesn't give you any protection against.

>you can train an AI on copyrighted material just like people can learn from copyrighted material. No amount of whining and hand wringing from engineers will ever make this true. This is for the courts to decide. A reasonable interpretation, in my eyes, is that the training process is a black box which takes in copyrighted works and produces a training model. The training model is a derivative work of the inputs. It…

Derivative works are typically things like language translations or film adaptions of a novel. A large language model is something like a probability breakdown of the order of word fragments in a body of text. It's a collection of statistics and math. It's different.

Now, can you get it to output a derivative work? Maybe. Is every output a derivative work? Maybe not.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#252
post #213
post #66

Earlier quoted context omitted.

How is it different from asking to me to summarize anything? I could have bought the book, or read the Wikipedia page, or listened people talking about it, or downloaded the torrent. In all those cases my summary could be right or could be wrong. If the rights holders know that I dowloaded the torrent they could sue me. In the other cases they can't. What if it turns out that OpenAI bought a copy of every book ingest…

> What if it turns out that OpenAI bought a copy of every book ingested be ChatGPT? That still doesn't necessarily confer to them the right to use it to train a model and generate derivative works based on purchased content.

[deleted]

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#253

Earlier quoted context omitted.

That's not was he's saying at all. He's saying you can train an AI on copyrighted material just like people can learn from copyrighted material. If you acquire the material illegally that a separate issue that training AI doesn't give you any protection against.

>you can train an AI on copyrighted material just like people can learn from copyrighted material. No amount of whining and hand wringing from engineers will ever make this true. This is for the courts to decide. A reasonable interpretation, in my eyes, is that the training process is a black box which takes in copyrighted works and produces a training model. The training model is a derivative work of the inputs. It…

> A reasonable interpretation, in my eyes, is that the training process is a black box which takes in copyrighted works and produces a training model. The training model is a derivative work of the inputs. It therefore violates the copyrights of a large number of rights holders. The outputs of the model are derivative works which also violate copyright.

What is the blackbox “limit” here? Is the mean value of all images in imagenet (which contains many copyrighted images) violating copyright? Is the character count of sarah silverman’s books? What about a prime number representing them - https://en.wikipedia.org/wiki/Illegal_number?wprov=sfti1

Training is much more similar to a character count than an illegal prime in my view, and thus, is almost certainly going to be okay/found to be okay. If not, something like, 90% of all models used today had some component trained on copyrighted data of some form.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#254

Earlier quoted context omitted.

He hacked into a server to release a database of paywalled studies to the public. Not only is it not the same but it was the hacking that brought charges upon him.

It's been quite a few years, but AFAIK he didn't hack JSTOR. He downloaded papers en mass using a guest MIT account that had legal access to JSTOR. Maybe he violated their terms, but that is not illegal. He did illegally trespass an unlocked MIT switch closet to do this. They blocked several IPs but his script would rotate to continue. The downloading was over a week or two, enough for security to set up a camera in…

> AFAIK he didn't hack JSTOR. He downloaded papers en mass using a guest MIT account that had legal access to JSTOR.

Potato, potahto. Or, like kids these days say it, "corporate wants you to find differences between these two pictures...".

Fact is, from the POV of the legal system, "using a guest account that had legal access to" a system, but to which (the account) you didn't have legal access, would typically be seen as hacking. So is running curl in a loop, if it results in you getting sued for it. So is just guessing the URL (e.g. incrementing a user ID in a GET query param), if it lets you access things you shouldn't be able to.

Yes, it's not aligned with how technology works. But it is aligned with expectations of behavior, which is what the law is really about.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#255

Earlier quoted context omitted.

Strictly speaking, it's uploading that people get sued for, not downloading. You can download all that you want from Z-Library or BitTorrent, as long as you don't share back. And indexing copyrighted material for search is safe, or at least ambiguous.

Downloading is illegal. That people do not normally get sued or prosecuted for downloading does not mean that they cannot get sued or prosecuted.

Is illegal where? Certainly not over here in Finland.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#256
post #224

Earlier quoted context omitted.

There's no "for thee but not for me" issue here: nobody has ever been sued or prosecuted simply for downloading, acquiring, or possessing illegally acquired copyrighted works. People are sued and prosecuted for unlicensed distribution . Making and having your own copies, and doing what you want with them, has always been fine. At worst it's a grey area, but in many cases it's been protected as fair use.

I am not sure where you're getting your information from, but you're not a lawyer, so you shouldn't be so confident in telling people what things are fine legally. Whether people have been sued for downloading works they don't have the right to copy onto their machines is irrelevant to whether it is actually illegal. And it certainly has nothing to do with fair use, which is about copyrighted works that you actually…

I've followed the issue in the US since the early 2000s as an activist and policy expert.

I'm not familiar with the state of play outside the US, but the US is one of the stricter jurisdictions in this regard, for reasons that have mostly to do with sophisticated corruption.

I'm responding to the "for me not for thee" and the top comment about there being an inconsistency between the treatment of large companies and the treatment of individuals in this case.

Unless people are typically punished for downloading and using copyrighted content, there is no such inconsistency. They are not, so there is not.

Copyright troll lawsuits have been fairly public and widely covered in the tech press, and most criminal prosecutions come with a formulaic, gloating press release from the law enforcement folks responsible. So it's pretty easy to follow this stuff.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#257
post #164

Earlier quoted context omitted.

Machine learning models have been trained with copyrighted data for a long time. Imagenet is full of copyrighted images, clearview literally just scanned the internet for faces, and I am sure there are other, older examples. I am unsure if this has been tested as fair use by a US court, but I am guessing it will be considered to be so if it is not already.

Excellent, so you're saying I'll be able to download any copyrighted work from any pirate site and be free of all consequence if I just claim that I'm training an AI?

> Excellent, so you're saying I'll be able to download any copyrighted work from any pirate site and be free of all consequence if I just claim that I'm training an AI?

Obviously not.

This kind of "So what you're saying is" exists to push the responders ideas, not the original speakers -- otherwise they wouldn't need to rephrase it so egregiously.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#258
post #176

Earlier quoted context omitted.

But acquiring the material illegally is the thrust of the issue, and the point of the thread here: the notion that large companies can get away with piracy if they just execute it on a large enough scale. Copyright infringment for thee but not for me.

Large countries seem to be doing it as well - https://petapixel.com/2023/06/05/japan-declares-ai-training-... And as long as OpenAI have an office in Japan they can absolutely legally train the models, no?

They could legally train the models in Japan, yes. Whether they could then use that model outside of Japan would be (and, possibly related, whether it would be considered a derivative work in those locales) would be ultimately up to the courts.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#259
post #164

Earlier quoted context omitted.

Excellent, so you're saying I'll be able to download any copyrighted work from any pirate site and be free of all consequence if I just claim that I'm training an AI?

You can currently download any copyright work from any pirate site and be free of all consequences, and this has always been so. You just can't upload, since that counts as distribution, triggering civil and criminal penalties written in an age before the Internet when only shady commercial operators would distribute unlicensed copyrighted works.

> You can currently download any copyright work from any pirate site and be free of all consequences, and this has always been so.

No.

By virtue of "download" of a file, you are making a copy of it which is in violation of US copyright (and lots of countries.

You're unlikely to be sued or prosecuted for it, but that doesn't make it legal.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#260

Earlier quoted context omitted.

In this case the correct analogy would be you brought a stolen painting into your house, looked at it for a while, and then produced your derivative work. Surely you see the issue here? Receiving stolen property?

Yes, I acknowledged that piracy is illegal in my previous post. That's not what the lawsuit is about, according to The Verge: >In the OpenAI suit, the trio offers exhibits showing that when prompted, ChatGPT will summarize their books, infringing on their copyrights.

That serves as evidence that the model has seen the material, and the only way the model could have seen the material is if it was pirated.
Post reply on HN