Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

121–130 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#121
post #92

Are we all reading the same complaint? They say: > in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private tracker.” Does that stack up? The Meta Paper -…

We don't seem to be reading the same thing, you're pulling Google out of thin air somewhere.

You might as well be complaining about the grammar. This is what was said in the article.

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private tracker.” Bibliotik and the other “shadow libraries” listed, says the lawsuit, are “flagrantly illegal.”

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#122
I am baffled by the fact that no enterprising lawyer so far have figured out the potential for a class action here.

Note: I am not telling whether I agree or not with such a class action, just pointing that it seems at least feasible and it could be potentially very lucrative for the lawyers involved. Of course, IANAL and all other disclaimers you can think of.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#123

Earlier quoted context omitted.

I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification. Aaron Swartz was a saint.

Swartz distributed information for everyone to use freely. These companies are processing it privately to develop their for-profit products. Big difference, IMO.

I am sad about closed source LLMs like ChatGPT, but Llama is in that grey area where it's freely available if you choose to ignore their silly license stuff, which of course pirates and AI developers are all too keen to ignore.

Even if they win the lawsuit, LLM development will simply go underground, and as we see from what the coomers at civitai and in the stable diffusion world have done, that may in fact ironically speed up development in AI.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#124
post #52

Getty Images also filed an AI lawsuit, alleging that Stability AI ... lol, bad karma? So it is okay for Getty to steal from others, but not ok for others to steal from them? I don't have a dog in this fight, but the goddamn the hypocrisy of these companies...

who does Getty steal from?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#125
post #90
post #36

Earlier quoted context omitted.

> Is there any legal basis for saying fair use permits distributing an LLM trained on copyrighted material, but you have to purchase all the content first to do so legally if it's only available for sale? My understanding (disclaimer: IANAL) is that in order to claim fair use, you have to be legally in possession of the work. If the work is only legally available for sale, then you must have legally purchased a copy,…

> you must have legally purchased a copy, or been given it by someone who did so (for example, if you received it as a gift). I am also NAL, but can I imagine it goes further than that. Just purchasing a copy doesn't let you create and sell (directly as content or indirectly via a service like a chatbot) derivative works that are substantially similar in style and voice to the original work. For example, an LLM 's re…

This is not how U.S. copyright law works. In order for something to eligible for copyright protection, it must be "fixed in a tangible medium of expression". Someone's exact words can be copyrighted, but their ideas or their style cannot be.

https://www.law.cornell.edu/wex/fixed_in_a_tangible_medium_o...

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#126
post #108

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

The lawsuit doesn't even mention Google.

No, I did. What's your point?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#127

Sarah's pov raises some questions for me regarding my own "training", there is a noteworthy part of who I am built upon the consumed music, books, movies, video games and tv shows that myself or people around me have pirated and shared with me. This part of me helped me in life appreciably, I could also say I profited because of it, helping me along my life in being likable, funny, relatable, with broad outlooks etc.…

Tom Scott did it: https://www.youtube.com/watch?v=IFe9wiDfb0E

The simple fact is that our current handling of copyright is just completely broken on so many levels.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#129
post #103

Earlier quoted context omitted.

We don't seem to be reading the same thing, you're pulling Google out of thin air somewhere.

I'm literally quoting The Verge article and following the links they present...

That paper is by Meta AI. Where are you getting Google from?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#130

This is actually quite interesting, as it's drawing a distinction between training material that can be accessed by anybody with a web browser (like anybody's blog), vs. training material that was "illegally-acquired... available in bulk via torrent systems." I don't think there's any reason why this would be a relevant legal distinction in terms of distributing an LLM -- blog authors weren't giving consent either. H…

Seems like the AI angle is just capitalizing on hype. If it's illegal to download "pirate" copyright material, that was the crime. The rest is basically irrelevant. If I watch a pirated movie, it's not illegal for me to tell someone the plot.
Post reply on HN