> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…
Sarah Silverman is suing OpenAI and Meta for copyright infringement
191–200 of 599 posts
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#192Earlier quoted context omitted.
I'm not entirely up to speed on US law, but wouldn't OpenAI have to provide the court some kind of proof that they didn't use it in the training data during discovery?
Not a layer, but I believe the plaintiff (the author) would need to prove that it regurgitates their copyrighted work - otherwise it is possibly fair use. OpenAI does not need to prove anything, just defend their position at a reasonable level. It’s not been decided if training a model on copyrighted works is “okay” or not as far as I know, but I expect it to be so, given that literally everyone does so at this point…
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#193Well theyre claiming their books were scraped illegally from torrents. if i torrent peter pan and watch it alone i can get thrown on jail. If AI is using torrents and getting billions in funding and revenue off a torrented peter pan they should probably be held to the same standard i am.
Now making copies and selling them on the street corner is another story.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#194Earlier quoted context omitted.
Excellent, so you're saying I'll be able to download any copyrighted work from any pirate site and be free of all consequence if I just claim that I'm training an AI?
That's not was he's saying at all. He's saying you can train an AI on copyrighted material just like people can learn from copyrighted material. If you acquire the material illegally that a separate issue that training AI doesn't give you any protection against.
No amount of whining and hand wringing from engineers will ever make this true. This is for the courts to decide.
A reasonable interpretation, in my eyes, is that the training process is a black box which takes in copyrighted works and produces a training model. The training model is a derivative work of the inputs. It therefore violates the copyrights of a large number of rights holders. The outputs of the model are derivative works which also violate copyright.
And anyone using or training a model trained on works for which they do not have the rights? Completely fucked. Or at least, they must accept this as a real risk.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#195Earlier quoted context omitted.
In Germany if you torrent stuff (without a VPN), you're very likely to get a letter from a law firm on behalf of the copyright holders saying that they'll sue you unless you pay them a nice flat fee of around 1000 Euro. It's no idle threat, and they will win if it goes to court.
That's because, when torrenting, you're typically also seeding a copy of it, i.e. you're distributing your local copy to other devices, and thus you're directly aiding in piracy. Simply downloading content from a centralized server, as explained above, is different. Although, one could argue what OpenAI & Meta are doing is closer to the torrent definition than the "simply downloading" definition, given that they're u…
This clearly needs some sort of regulation or policy.
It's clearly pretty bullshit if you ask chatgpt for a joke and it repeats a Sarah Silverman joke to you, while they charge you a subscription for it and she gets none of that sub money.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#196I am baffled by the fact that no enterprising lawyer so far have figured out the potential for a class action here. Note: I am not telling whether I agree or not with such a class action, just pointing that it seems at least feasible and it could be potentially very lucrative for the lawyers involved. Of course, IANAL and all other disclaimers you can think of.
Looking at the two PDFs embedded at the bottom of the The Verge article, they both say "class action" on the first page, and they say the three plaintiffs are suing "on behalf of themselves and all other similarly situated".
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#197Earlier quoted context omitted.
That's not was he's saying at all. He's saying you can train an AI on copyrighted material just like people can learn from copyrighted material. If you acquire the material illegally that a separate issue that training AI doesn't give you any protection against.
But acquiring the material illegally is the thrust of the issue, and the point of the thread here: the notion that large companies can get away with piracy if they just execute it on a large enough scale. Copyright infringment for thee but not for me.
Also, the notion of the downloading itself being an illegal act is not universal as others have pointed out.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#198Earlier quoted context omitted.
> If you downloaded a book from that website, you would be sued and found guilty of infringement. How often does this actually happen? You might get handed an infringement notice, and your ISP might terminate your service if you're really egregious about it, but I haven't ever heard of someone actually being sued for downloading something.
>How often does this actually happen? Did you hear about Aaron Schwartz?
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#199Earlier quoted context omitted.
If you found a way to have a million children who could grow up in one day your analogy would be more apt. In that case you and your children would rightly be considered a threat.
did you read to the third analogy?
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#200>On information and belief, the reason ChatGPT can accurately summarize a certain copyrighted book is because that book was copied by OpenAI and ingested by the underlying OpenAI Language Model (either GPT-3.5 or GPT-4) as part of its training data. While it strikes me as perfectly plausible that the Books2 dataset contains Silverman's book, this quote from the complaint seems obviously false. First, even if the mode…
Plausible gets you discovery. Discovery gets you closer to the what the actual facts are.