Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

131–140 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#131

Earlier quoted context omitted.

Isn’t the burden of proof on the other side?

Not when OpenAI publicly declared they trained on pirated works. I can’t imagine “we can’t tell if this is the result of the illegal thing we did or not” is going to stand up very well, nor does it bode well for any refutation of the plaintiff’s depiction of their intent. Part of fair use consideration is commercial impact and when you steal a bunch of books to train your AI model, it’s hard to refute that the impact…

Please read more carefully. OpenAI never “declared they trained on pirated works.”

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#133
post #90
post #36

Earlier quoted context omitted.

> Is there any legal basis for saying fair use permits distributing an LLM trained on copyrighted material, but you have to purchase all the content first to do so legally if it's only available for sale? My understanding (disclaimer: IANAL) is that in order to claim fair use, you have to be legally in possession of the work. If the work is only legally available for sale, then you must have legally purchased a copy,…

> you must have legally purchased a copy, or been given it by someone who did so (for example, if you received it as a gift). I am also NAL, but can I imagine it goes further than that. Just purchasing a copy doesn't let you create and sell (directly as content or indirectly via a service like a chatbot) derivative works that are substantially similar in style and voice to the original work. For example, an LLM 's re…

> ... IMO doesn't constitute fair use, because the intellectual property of the artist is their style even more than the content they produce

You're essentially banning satire here, though. There's plenty of folks making a living as cover bands or impersonators. I'm not sure what the answer is, but it's definitely not outright outlawing imitation.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#134

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification. Aaron Swartz was a saint.

While I’ma proponent of free information and loosening copyright, allowing billion dollar companies to package up the sum of human creation and resell statistical models that mimic the content and style of everyone… is a bit far.

Fair use is for humans.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#135

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification. Aaron Swartz was a saint.

Copyrights shouldn't exist, but

> LLM weights, and the datasets are "fair-use" or whatever other silly legal justification.

Would just be a carve out for the wealthy. If these laws don't mean anything, everyone who got harassed, threatened, extorted, fined, arrested, tried, or jailed for internet piracy are owed reparations. Let Meta pay them.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#136

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification. Aaron Swartz was a saint.

[dead]

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#137

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification. Aaron Swartz was a saint.

> I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification.

I would be very happy if either a court or lawmakers decided that copyright itself was unconscionable. That isn't what's going to happen, though. And I think it's incredibly unacceptable if a court or lawmakers instead decide that AI training in particular gets a special exception to violate other people's copyrights on a massive scale when nobody else gets to do that.

As far as a fair use argument in particular, fair use in the US is a fourfold test:

> the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes;

The purpose and character is absolutely heavily commercial and makes a great deal of money for the companies building the AIs. A primary use is to create other works of commercial value competing with the original works.

> the nature of the copyrighted work;

There's nothing about the works used for AI training that makes them any less entitled to copyright protections or more permissive of fair use than anything else. They're not unpublished, they're not merely collections of facts or ideas, etc.

> the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and

AI training uses entire works, not excerpts.

> the effect of the use upon the potential market for or value of the copyrighted work.

AI models are having a massive effect on the market for and value of the works they train on, as is being widely discussed in multiple industries. Art, writing, voice acting, code; in any market AI can generate content for, the value of such content goes down. (This argument does not require AI to be as good as humans. Even if the high end of every market produces work substantially better than AI, flooding a market with unlimited amounts of cheap/free low-end content still has a massive effect.)

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#138
post #79
post #66

Earlier quoted context omitted.

How is it different from asking to me to summarize anything? I could have bought the book, or read the Wikipedia page, or listened people talking about it, or downloaded the torrent. In all those cases my summary could be right or could be wrong. If the rights holders know that I dowloaded the torrent they could sue me. In the other cases they can't. What if it turns out that OpenAI bought a copy of every book ingest…

I feel like thats one of the many questions regulators and law makers are going to be asked long term. I'm sure buying the book for "commercial purposes" like that would't be appropriate, but then again, does that mean if I read it and then summarize it in my work, or regurgitate its info as part of my job...I'm violating a license? A world where humans have special permissions but LLMs don't seems pretty interesting…

For "LLMs" read "corporations" (it's not the LLM trying to argue that copyright applies to you but not them) and this seems... possibly ok?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#139
post #102

Earlier quoted context omitted.

Making an analogy where you substitute a human being for the LLM is disingenuous to the extreme.

why do you think that?

Because LLMs are not people. They are nothing like people; not in construction nor behaviour.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#140

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification. Aaron Swartz was a saint.

There's a difference between "information wants to be free" and "Facebook can produce works minimally derived from your greatest creative work at a scale you can't match". LLMs seem to aggregate that value to whoever builds the model, which they can then sell access to, or sell the output it produces.

Five years from now, will OpenAI actually be open, or will it be a rent seeking org chasing the next quarterly gains? I expect the latter.

Post reply on HN