Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

191–200 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#191

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

[deleted]

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#192

Earlier quoted context omitted.

I'm not entirely up to speed on US law, but wouldn't OpenAI have to provide the court some kind of proof that they didn't use it in the training data during discovery?

Not a layer, but I believe the plaintiff (the author) would need to prove that it regurgitates their copyrighted work - otherwise it is possibly fair use. OpenAI does not need to prove anything, just defend their position at a reasonable level. It’s not been decided if training a model on copyrighted works is “okay” or not as far as I know, but I expect it to be so, given that literally everyone does so at this point…

No. Fair use is an affirmative defense to infringement. If they admit to infringement or in the alternative want to argue fair use, the burden is on OpenAI to demonstrate their use falls within the relevant standard for fair use.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#193

Well theyre claiming their books were scraped illegally from torrents. if i torrent peter pan and watch it alone i can get thrown on jail. If AI is using torrents and getting billions in funding and revenue off a torrented peter pan they should probably be held to the same standard i am.

I don't think you can be thrown in jail. Just torrenting a film and watching it would be a civil offense?

Now making copies and selling them on the street corner is another story.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#194
post #164

Earlier quoted context omitted.

Excellent, so you're saying I'll be able to download any copyrighted work from any pirate site and be free of all consequence if I just claim that I'm training an AI?

That's not was he's saying at all. He's saying you can train an AI on copyrighted material just like people can learn from copyrighted material. If you acquire the material illegally that a separate issue that training AI doesn't give you any protection against.

>you can train an AI on copyrighted material just like people can learn from copyrighted material.

No amount of whining and hand wringing from engineers will ever make this true. This is for the courts to decide.

A reasonable interpretation, in my eyes, is that the training process is a black box which takes in copyrighted works and produces a training model. The training model is a derivative work of the inputs. It therefore violates the copyrights of a large number of rights holders. The outputs of the model are derivative works which also violate copyright.

And anyone using or training a model trained on works for which they do not have the rights? Completely fucked. Or at least, they must accept this as a real risk.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#195
post #150
post #141

Earlier quoted context omitted.

In Germany if you torrent stuff (without a VPN), you're very likely to get a letter from a law firm on behalf of the copyright holders saying that they'll sue you unless you pay them a nice flat fee of around 1000 Euro. It's no idle threat, and they will win if it goes to court.

That's because, when torrenting, you're typically also seeding a copy of it, i.e. you're distributing your local copy to other devices, and thus you're directly aiding in piracy. Simply downloading content from a centralized server, as explained above, is different. Although, one could argue what OpenAI & Meta are doing is closer to the torrent definition than the "simply downloading" definition, given that they're u…

Honestly don't think our current laws are even good for this case.

This clearly needs some sort of regulation or policy.

It's clearly pretty bullshit if you ask chatgpt for a joke and it repeats a Sarah Silverman joke to you, while they charge you a subscription for it and she gets none of that sub money.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#196

I am baffled by the fact that no enterprising lawyer so far have figured out the potential for a class action here. Note: I am not telling whether I agree or not with such a class action, just pointing that it seems at least feasible and it could be potentially very lucrative for the lawyers involved. Of course, IANAL and all other disclaimers you can think of.

These lawsuits are class action suits, aren't they?

Looking at the two PDFs embedded at the bottom of the The Verge article, they both say "class action" on the first page, and they say the three plaintiffs are suing "on behalf of themselves and all other similarly situated".

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#197
post #176

Earlier quoted context omitted.

That's not was he's saying at all. He's saying you can train an AI on copyrighted material just like people can learn from copyrighted material. If you acquire the material illegally that a separate issue that training AI doesn't give you any protection against.

But acquiring the material illegally is the thrust of the issue, and the point of the thread here: the notion that large companies can get away with piracy if they just execute it on a large enough scale. Copyright infringment for thee but not for me.

The power imbalance has always been there. Getty images commits large amounts of copyright theft and mostly gets away with it, but will happily sue the shit out of you for using your own images they stole.

Also, the notion of the downloading itself being an illegal act is not universal as others have pointed out.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#198
post #181

Earlier quoted context omitted.

> If you downloaded a book from that website, you would be sued and found guilty of infringement. How often does this actually happen? You might get handed an infringement notice, and your ISP might terminate your service if you're really egregious about it, but I haven't ever heard of someone actually being sued for downloading something.

>How often does this actually happen? Did you hear about Aaron Schwartz?

He hacked into a server to release a database of paywalled studies to the public. Not only is it not the same but it was the hacking that brought charges upon him.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#199
post #78

Earlier quoted context omitted.

If you found a way to have a million children who could grow up in one day your analogy would be more apt. In that case you and your children would rightly be considered a threat.

did you read to the third analogy?

What if in your third analogy you replace millions of students by billions of processes running on machines, each of which can generate output ten thousand times faster than a college educated human?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#200
post #38

>On information and belief, the reason ChatGPT can accurately summarize a certain copyrighted book is because that book was copied by OpenAI and ingested by the underlying OpenAI Language Model (either GPT-3.5 or GPT-4) as part of its training data. While it strikes me as perfectly plausible that the Books2 dataset contains Silverman's book, this quote from the complaint seems obviously false. First, even if the mode…

Plausible is literally the standard to clear a motion to dismiss.

Plausible gets you discovery. Discovery gets you closer to the what the actual facts are.

Post reply on HN