Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

201–210 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#201

Earlier quoted context omitted.

> I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification. I would be very happy if either a court or lawmakers decided that copyright itself was unconscionable. That isn't what's going to happen, though. And I think it's incredibly unacceptable if a court…

It is the equivalent of making a 3D map of a museum and getting sued by one artist of one painting in the museum. Ant individual work in an AI dataset is nearly worthless - only in aggregate does it have value. If that doesn't count as a "transformative work" I don't know what does.

They could literally just repeat a Silverman routine verbatim

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#202

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Strictly speaking, it's uploading that people get sued for, not downloading.

You can download all that you want from Z-Library or BitTorrent, as long as you don't share back. And indexing copyrighted material for search is safe, or at least ambiguous.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#203
post #164

Earlier quoted context omitted.

Machine learning models have been trained with copyrighted data for a long time. Imagenet is full of copyrighted images, clearview literally just scanned the internet for faces, and I am sure there are other, older examples. I am unsure if this has been tested as fair use by a US court, but I am guessing it will be considered to be so if it is not already.

Excellent, so you're saying I'll be able to download any copyrighted work from any pirate site and be free of all consequence if I just claim that I'm training an AI?

You can currently download any copyright work from any pirate site and be free of all consequences, and this has always been so.

You just can't upload, since that counts as distribution, triggering civil and criminal penalties written in an age before the Internet when only shady commercial operators would distribute unlicensed copyrighted works.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#204

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Machine learning models have been trained with copyrighted data for a long time. Imagenet is full of copyrighted images, clearview literally just scanned the internet for faces, and I am sure there are other, older examples. I am unsure if this has been tested as fair use by a US court, but I am guessing it will be considered to be so if it is not already.

> I am unsure if this has been tested as fair use by a US court

Not yet. One suit that a lot of us are watching is the GitHub co-pilot lawsuit: https://githubcopilotlitigation.com/

There is a prediction market for it, currently trading at 19%: https://manifold.markets/JeffKaufman/will-the-github-copilot...

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#205

Sarah's pov raises some questions for me regarding my own "training", there is a noteworthy part of who I am built upon the consumed music, books, movies, video games and tv shows that myself or people around me have pirated and shared with me. This part of me helped me in life appreciably, I could also say I profited because of it, helping me along my life in being likable, funny, relatable, with broad outlooks etc.…

Just because we call both learning, doesn't mean that human learning and machine learning are the same. They most definitely are not the same. Human learning is very lossy.

Even if they were the same, it doesn't mean that bots should have the same rights that people have.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#206
post #50

This is actually quite interesting, as it's drawing a distinction between training material that can be accessed by anybody with a web browser (like anybody's blog), vs. training material that was "illegally-acquired... available in bulk via torrent systems." I don't think there's any reason why this would be a relevant legal distinction in terms of distributing an LLM -- blog authors weren't giving consent either. H…

Eventually, I imagine a new licensing concept will emerge, similar to the idea of music synchronization rights -- maybe call it "training rights." It won't matter whether the text was purchased or pirated -- just like it doesn't matter now if an audio track was purchased or pirated, when it's mixed into in a movie soundtrack. Talent agencies will negotiate training rights fees in bulk for popular content creators, wh…

Humans are also trained on copyrighted content they see. Should every artist have to pay that fee too on every work they create?

Disney will finally be able to charge a "you know what the mouse looks like" tax.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#207

This is actually quite interesting, as it's drawing a distinction between training material that can be accessed by anybody with a web browser (like anybody's blog), vs. training material that was "illegally-acquired... available in bulk via torrent systems." I don't think there's any reason why this would be a relevant legal distinction in terms of distributing an LLM -- blog authors weren't giving consent either. H…

The "fun" part about cases like this is that we don't really know what the contours of the law are as applied to training data like this. Illegally downloading a book is an independent act of infringement (to my recollection at least). So I'm not sure that it matters if you eventually trained an LLM with it vs read for your own enjoyment. But we will see! Fair use is a possibility here but we need a court to apply the test and that will probably go up to SCOTUS eventually.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#208
post #176

Earlier quoted context omitted.

That's not was he's saying at all. He's saying you can train an AI on copyrighted material just like people can learn from copyrighted material. If you acquire the material illegally that a separate issue that training AI doesn't give you any protection against.

But acquiring the material illegally is the thrust of the issue, and the point of the thread here: the notion that large companies can get away with piracy if they just execute it on a large enough scale. Copyright infringment for thee but not for me.

There's no "for thee but not for me" issue here: nobody has ever been sued or prosecuted simply for downloading, acquiring, or possessing illegally acquired copyrighted works. People are sued and prosecuted for unlicensed distribution.

Making and having your own copies, and doing what you want with them, has always been fine. At worst it's a grey area, but in many cases it's been protected as fair use.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#209

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Strictly speaking, it's uploading that people get sued for, not downloading. You can download all that you want from Z-Library or BitTorrent, as long as you don't share back. And indexing copyrighted material for search is safe, or at least ambiguous.

Carefully speaking, what you say is true in many places (countries), but also not true in other places (countries). Some jurisdictions are different, as always.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#210

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Strictly speaking, it's uploading that people get sued for, not downloading. You can download all that you want from Z-Library or BitTorrent, as long as you don't share back. And indexing copyrighted material for search is safe, or at least ambiguous.

Huh? No.
Post reply on HN