Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

141–150 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#141

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

> If you downloaded a book from that website, you would be sued and found guilty of infringement. How often does this actually happen? You might get handed an infringement notice, and your ISP might terminate your service if you're really egregious about it, but I haven't ever heard of someone actually being sued for downloading something.

In Germany if you torrent stuff (without a VPN), you're very likely to get a letter from a law firm on behalf of the copyright holders saying that they'll sue you unless you pay them a nice flat fee of around 1000 Euro.

It's no idle threat, and they will win if it goes to court.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#142

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

> If you downloaded a book from that website, you would be sued and found guilty of infringement. How often does this actually happen? You might get handed an infringement notice, and your ISP might terminate your service if you're really egregious about it, but I haven't ever heard of someone actually being sued for downloading something.

If books aren't under copyright protection and they're entirely legal to download, I agree that this lawsuit has no merit.

If that's not what you're saying, I don't understand your point. Is it the difference between the phrases "would be" and "could be," or even "should be"?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#143
post #139

Earlier quoted context omitted.

why do you think that?

Because LLMs are not people. They are nothing like people; not in construction nor behaviour.

LLMs are built upon neural networks which are modelled upon how brains work

can you explain to me specifically how they're different?

can you explain to me how they're different to the degree that making an analogy between the two is "disingenuous to the extreme"?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#144

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

> If you downloaded a book from that website, you would be sued and found guilty of infringement. How often does this actually happen? You might get handed an infringement notice, and your ISP might terminate your service if you're really egregious about it, but I haven't ever heard of someone actually being sued for downloading something.

Exactly, never happens. It's a threat parents and teachers tell school children to try to spook them from pirating but it isn't financially worth it for an author or publishing company to try to sue an individual over some books or music downloads. The only cases are business to business over mass downloads where it could make financial sense to pay for lawyers to sue.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#145
post #50

Earlier quoted context omitted.

Eventually, I imagine a new licensing concept will emerge, similar to the idea of music synchronization rights -- maybe call it "training rights." It won't matter whether the text was purchased or pirated -- just like it doesn't matter now if an audio track was purchased or pirated, when it's mixed into in a movie soundtrack. Talent agencies will negotiate training rights fees in bulk for popular content creators, wh…

>> Talent agencies will negotiate training rights fees in bulk for popular content creators AFAICT there is no legal recognition of "training rights" or anything similar. First sale right is a thing, but even textbooks don't get extra rights for their training or educational value.

Many legal concepts used by courts has no legal recognition in the law texts. Much of legal practice are just precedents, policies, customs, and doctrines.

Parent comment mention music synchronization rights, and this concept does not exist in copyright. Court do occasionally mention it, and lawyers talks about it, but in terms of the legal recognition there is basically only the law text that define derivative work and fair use. One way to interpret it is that court has precedents to treat music synchronization as a derivative work that do not fall under fair use.

Using textbooks in training/education is not as black and white that one may assume. Take this Berkeley (https://teaching.berkeley.edu/resources/course-design/using-...). Copying in this context include using pages for slides and during lectures (which is a slightly large scope than making physical copies on physical paper). In obvious case the answer is likely obvious, but in others it will be more complex.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#146
post #139

Earlier quoted context omitted.

Because LLMs are not people. They are nothing like people; not in construction nor behaviour.

LLMs are built upon neural networks which are modelled upon how brains work can you explain to me specifically how they're different? can you explain to me how they're different to the degree that making an analogy between the two is "disingenuous to the extreme"?

One is a person, the other is a computer program.

Legally quite distinct! Note that nobody is even seriously claiming we have an AGI, there's no Star Trek discussion of whether an android is a person. Everyone agrees this is just a computer program.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#147

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

> If you downloaded a book from that website, you would be sued and found guilty of infringement. How often does this actually happen? You might get handed an infringement notice, and your ISP might terminate your service if you're really egregious about it, but I haven't ever heard of someone actually being sued for downloading something.

Whether or not it's enforced, it's illegal and copyright holders are within their rights to sue you. This is piratebay levels of piracy but because it's done by a large company and is sufficiently obfuscated behind tech, people don't see it the same way.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#148

Earlier quoted context omitted.

I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification. Aaron Swartz was a saint.

While I’ma proponent of free information and loosening copyright, allowing billion dollar companies to package up the sum of human creation and resell statistical models that mimic the content and style of everyone… is a bit far. Fair use is for humans.

Yeah, but hypothetically should open source projects be offered special protections? I feel like they should, and with certain caveats where, say, a company like Meta is allowed to claim fair use if and only if they free up the entire ecosystem as they deploy it.

But yeah, having private ostensibly profitable models based on other people's work without giving them free access to it is not fair. Give some get some.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#149
post #139

Earlier quoted context omitted.

Because LLMs are not people. They are nothing like people; not in construction nor behaviour.

LLMs are built upon neural networks which are modelled upon how brains work can you explain to me specifically how they're different? can you explain to me how they're different to the degree that making an analogy between the two is "disingenuous to the extreme"?

The differences between a a human being and a computer are too numerous to list. I don’t even know why you need to ask the question.

Let me ask another question to point out the absurdity of yours: Human beings have more in common with a bacterium than a software program. Can you tell me specifically how humans are not bacteria?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#150
post #141

Earlier quoted context omitted.

> If you downloaded a book from that website, you would be sued and found guilty of infringement. How often does this actually happen? You might get handed an infringement notice, and your ISP might terminate your service if you're really egregious about it, but I haven't ever heard of someone actually being sued for downloading something.

In Germany if you torrent stuff (without a VPN), you're very likely to get a letter from a law firm on behalf of the copyright holders saying that they'll sue you unless you pay them a nice flat fee of around 1000 Euro. It's no idle threat, and they will win if it goes to court.

That's because, when torrenting, you're typically also seeding a copy of it, i.e. you're distributing your local copy to other devices, and thus you're directly aiding in piracy. Simply downloading content from a centralized server, as explained above, is different.

Although, one could argue what OpenAI & Meta are doing is closer to the torrent definition than the "simply downloading" definition, given that they're using that to redistribute information to others. It'll be an interesting case.

Post reply on HN