Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

351–360 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#351

Earlier quoted context omitted.

There are already many more public domain books than people are inclined to read: https://www.gutenberg.org/ https://librivox.org/ many of which form the basis for an education: https://news.ycombinator.com/item?id=34630153 And most of which, when in copyright, paid their authors quite handsomely in terms of royalties. If you believe that books should exist without copyright, then one has to ask --- how many books ha…

I’d more ask what the cost vs benefits are of keeping the existing scheme, it’s not free to run all these DRM services, prosecute offenders etc… Not to say I support no copyright…

The benefit of the existing scheme is that new works get created, and some of them are even copy-edited and published.

How many texts are created which are explicitly placed in the public domain and from which the authors have made a conscious decision not to profit thereby?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#352

> The lawsuit against OpenAI alleges that summaries of the plaintiffs’ work generated by ChatGPT indicate the bot was trained on their copyrighted content. “The summaries get some details wrong” but still show that ChatGPT “retains knowledge of particular works in the training dataset," the lawsuit says. Setting aside the whole issue of whether LLM constitutes a derived work of whatever it's trained on, this sounds l…

That isn't firm evidence, but courts don't need firm evidence to start a case and discover new facts. They very well can ask LLM experts, and openAI themselves, whether that output is highly likely to have been derived from the copyrighted work in question. Anyway. If the argument is "No, it's not from the book, it's from someone else's copyrighted summary", that just means the person who wrote such a summary needs t…

> that just means the person who wrote such a summary needs to instead sue for copyright infringement right?

Doesn't need to be a person, could be another AI that wrote the summaries. I see a big problem for copyrights looming on the horizon - LLMs can reword, rewrite or generate input-output pairs using copyrighted data as reference, thus creating clean data for training. AI cleanly separates knowledge from expression. And maybe it should do so just to reduce inconsistencies and PII in organic text.

Copyrights should only be concerned with expression not knowledge, right? Protecting knowledge is the object of patents, and protecting names the object of trademarks. Copyright is only related to expression otherwise it would become too powerful. For example, instead of banning reproduction of this paragraph, it would also cover all its possible paraphrases. That would be like owning an idea, the "*" version, not a unique sequence of words.

Does it even make sense to talk about copyrights when everything can be remade in many ways so easily? Copyright was already suffering greatly since zero cost copying became a thing, now LLMs are dealing the second blow. It's just a fig leaf by now.

If we take a step back, it's all knowledge and language, self replicating memes under an evolutionary force. It's language evolution, or idea evolution. We are just supporting it by acting as language agents, but now LLMs got into the game, so ideas got a new vector of self replication. We want to own this process piece by piece but such a thing might be arrogant and go against the trend. Knowledge wants to be free, it wants to mix and match, travel and evolve. This process looks like biology, it has a will of its own.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#353

I think this will be a bigger issue than some people think. Maybe there's a market for 'clean' training data that doesn't include potential copyright claims. Just public domain works. We'll know it's an AI because it talks like a late 18th century/early 19th century writer?

I think the onus should be on Sarah Silverman to prove which inputs were hers and which outputs leveraged those inputs.

I think she should pay all of the court costs if she fails to do so.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#354

Earlier quoted context omitted.

Let's take a second to remember that copyright is the reason ~every child doesn't have access to ~every book ever written. While it might be too disruptive to eliminate copyright overnight, we should remember that our world will be much better and improve much faster to the extent we can reduce copyright's impact. And we should cheer it on when it happens. A majority of the world's population in 2023 has a smartphone…

It would look largely identical to ours, I think. It's pretty trivial to get access to many, if not most, e-books. Any public-domain work is available on Project Gutenberg [0]. Copyrighted works can be accessed for free using tools of various legality: Libby [1] is likely sponsored by your local library and gives free access to e-books and audiobooks. Library Genesis [2] has a questionable legal status but has a huge…

Hoopla.com as well

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#355

Earlier quoted context omitted.

There's an interesting nuance here if you were to put a human in the place of the LLM. We have read thousands of works; does that mean anything we write is derivative?

Humans are special and can create new copyrights. The process of a human brain synthesizing stuff does act as a barrier to copyright infringement. Machines and algorithms are not legally recognized as being able to author original non-derivative works. > put a human in the place of the LLM But also, no, if you have a team of humans doing rote matrix multiplication instead of an LLM, that does not make it so the matri…

Not really. The "special" quality attached to humans is only in creating copyright -- it has nothing to do with fair use arguments around derivative works.

"Machines and algorithms are not legally recognized as being able to author original non-derivative works"

Neither are monkeys. This doesn't mean a monkey's painting is any more or less derivative, or any more or less subject to a copyright claim. It only means that there is not a second copyright attached to the resulting work.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#356

Earlier quoted context omitted.

This makes even less sense than the previous guy. When you've figured out how the world will work without people incentivized by money and power, get back to the rest of us

Suing people for reading a book is one of the ways the poor are kept poor. The world will work just fine without ceding control to people who seek money and power, because most people aren't like that. The question is how to prevent the few who are from oppressing the rest of us. That is quite a challenge, but haven't you ever created something just for the fun of it?

> Suing people for reading a book is one of the ways the poor are kept poor.

Who is suing anyone for reading a book?

> The world will work just fine without ceding control to people who seek money and power, because most people aren't like that. The question is how to prevent the few who are from oppressing the rest of us.

Please get me in contact with your dealer because apparently the “legal” stuff I’ve been getting is not as potent as I thought.

> That is quite a challenge, but haven't you ever created something just for the fun of it?

Sure I’ve painted things that are on my wall and created lots of utility apps that I use personally but I keep them to myself. If I thought they were something that could get me some extra pocket money or better then I would be looking for ways to monetize. I have zero interest in sharing my potential intellectual property for free.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#357

Earlier quoted context omitted.

Possession of copyrighted material without permission of the creator is illegal, but yes -- rights holders don't really go after infringers except maybe via ISP three strikes crap. They're very much incentivized to change their behavior for AI scraping, though.

I am not a lawyer, but my understanding is that copyright law typically regulates the unauthorized reproduction, distribution, public display, or creation of derivative works of copyrighted materials. Possession of copyrighted material in itself is not illegal. It's how you use that material that could potentially violate copyright laws.

> Possession of copyrighted material in itself is not illegal.

The means of procurement matters. If they are in possession of copyrighted material because someone without the proper rights gave it to them illegally, then the possession itself is also illegal. It's illegal to own knowingly stolen property in all 50 US states and most countries, and while we could argue to the end of days about whether copying a file truly qualifies as stealing, the legal precedents are very clear on the matter.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#358

Earlier quoted context omitted.

That's not was he's saying at all. He's saying you can train an AI on copyrighted material just like people can learn from copyrighted material. If you acquire the material illegally that a separate issue that training AI doesn't give you any protection against.

>you can train an AI on copyrighted material just like people can learn from copyrighted material. No amount of whining and hand wringing from engineers will ever make this true. This is for the courts to decide. A reasonable interpretation, in my eyes, is that the training process is a black box which takes in copyrighted works and produces a training model. The training model is a derivative work of the inputs. It…

> A reasonable interpretation, in my eyes, is that the training process is a black box which takes in copyrighted works and produces a training model. The training model is a derivative work of the inputs.

Its unmistakably not a derivative work of the inputs individually or collectively, since a derivative work must be itself an distinct work of authorship (the same as the work of authorship requirement for copyright), and the output of a purely mechanical process is not.

The collection of inputs itself might be a derivative work of the individual inputs, before considering Fair Use.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#359
post #344

Earlier quoted context omitted.

That's a very strong claim that for which you should probably provide some evidence. I just asked ChatGPT to produce a script of a scene from a movie by asking it first to change a single line (which wouldn't be transformative) and then asking it to restore the line to the original. It obliged. Sure, it's probably not the same exact script, but it's not transformative at all. In any case, the issue here isn't whether…

Okay, so you used a tool to duplicate a copyrighted work. You could do the same thing with a word processor. YOU are obviously the one who violated copyright by using the tool that way. I don't understand how a reasonable person could have a different interpretation.

OP's claim requires the AI to produce work. If you're attributing the work the AI produced as my work, that's fine; but it means you believe the AI doesn't produce work at all; which would render the OP's claim not only false, but nonsensical.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#360
post #224

Earlier quoted context omitted.

I am not sure where you're getting your information from, but you're not a lawyer, so you shouldn't be so confident in telling people what things are fine legally. Whether people have been sued for downloading works they don't have the right to copy onto their machines is irrelevant to whether it is actually illegal. And it certainly has nothing to do with fair use, which is about copyrighted works that you actually…

I've followed the issue in the US since the early 2000s as an activist and policy expert. I'm not familiar with the state of play outside the US, but the US is one of the stricter jurisdictions in this regard, for reasons that have mostly to do with sophisticated corruption. I'm responding to the "for me not for thee" and the top comment about there being an inconsistency between the treatment of large companies and…

Are you forgetting all the DMCA lawsuits slapping individuals who downloaded MP3s with tens of thousands of dollars? These were not corporations, these were teenagers still living with their parents who pulled music files off the likes of Napster.

The DMCA does allow harassment by copyright holders to individuals suspected of infringement. It's just that most people like authors wouldn't blow their legal budget suing kids.

Post reply on HN