Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

151–160 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#151
post #146

Earlier quoted context omitted.

LLMs are built upon neural networks which are modelled upon how brains work can you explain to me specifically how they're different? can you explain to me how they're different to the degree that making an analogy between the two is "disingenuous to the extreme"?

One is a person, the other is a computer program. Legally quite distinct! Note that nobody is even seriously claiming we have an AGI, there's no Star Trek discussion of whether an android is a person. Everyone agrees this is just a computer program.

are you aware of how neural networks work? and remember that this is simply an illustrative analogy

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#152
post #128

The summary and other LLM outputs should fall into transformative use. I wonder how they intend to prove as this use case is little different from a person reading the book and writing about it.

This site would be better if you got banned for commenting without reading the article.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#153
post #146

Earlier quoted context omitted.

LLMs are built upon neural networks which are modelled upon how brains work can you explain to me specifically how they're different? can you explain to me how they're different to the degree that making an analogy between the two is "disingenuous to the extreme"?

One is a person, the other is a computer program. Legally quite distinct! Note that nobody is even seriously claiming we have an AGI, there's no Star Trek discussion of whether an android is a person. Everyone agrees this is just a computer program.

It doesn't matter if it's a person, or a computer program, or not. This discussion is moot. Is there a substantial reproduction of the works in the output? If not, there's no copyright infringement here.

Try reading this legal opinion: https://lawreview.law.ucdavis.edu/issues/53/5/notes/files/53...

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#154
post #133
post #90

Earlier quoted context omitted.

> you must have legally purchased a copy, or been given it by someone who did so (for example, if you received it as a gift). I am also NAL, but can I imagine it goes further than that. Just purchasing a copy doesn't let you create and sell (directly as content or indirectly via a service like a chatbot) derivative works that are substantially similar in style and voice to the original work. For example, an LLM 's re…

> ... IMO doesn't constitute fair use, because the intellectual property of the artist is their style even more than the content they produce You're essentially banning satire here, though. There's plenty of folks making a living as cover bands or impersonators. I'm not sure what the answer is, but it's definitely not outright outlawing imitation.

> You're essentially banning satire here, though.

I specifically noted that I'm talking about limiting the rights of machine generated mimicry. Satire by a person is completely different and involves the satirist's own style and experience that is derived from their human experience. Alec Baldwin's Trump impersonation is quite different than Trevor Noah's, for example. I presume both were also written by people, not LLMs.

I fully support the satirical impersonation of politicians and celebrities, but I feel far less comfortable with LLM generated content in the style of Trump or Obama, especially when presented using voice synthesis, even when it is fully disclaimed as a fake.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#155

>> In the OpenAI suit, the trio offers exhibits showing that when prompted, ChatGPT will summarize their books, infringing on their copyrights. This doesn't seem like copyright infringement. I could read the book and offer a summary right? Someone on goodreads could as well. Why should an AI doing it be different? BTW I could also read someone's illicit copy and do the same, couldn't I? I think people are trying to c…

I think the argument is somewhat more interesting if the book was pirated --both you as an individual and OpenAI as a company could be sued for that. But I really don't see how you could prove OpenAI did that, since ChatGPT could have learned from existing summaries on Wikipedia and Goodreads.

It seems pretty easy to prove that, since they admitted it in public.

Read the article. This isn't about the question of LLMs being copyright infringement, this is about Meta and OpenAI admitting that they had pirated copies of those books.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#156
post #52

Getty Images also filed an AI lawsuit, alleging that Stability AI ... lol, bad karma? So it is okay for Getty to steal from others, but not ok for others to steal from them? I don't have a dog in this fight, but the goddamn the hypocrisy of these companies...

who does Getty steal from?

https://www.dpreview.com/news/3907450005/getty-images-sued-o...

> CixxFive Concepts, a digital marketing company based in Dallas, Texas, has filed a class action lawsuit against Getty Images over its alleged licensing of public domain images.

> Though CixxFive acknowledges that it is not illegal to sell public domain images, the company alleges that Getty's 'conduct goes much further than this,' claiming it has utilized 'a number of different deceptive techniques' in order to 'mislead' its customers -- and potential future customers -- into thinking the company owns the copyrights of all images it sells.

> The alleged actions, the lawsuit claims, 'purport to restrict the use of the public domain images to a limited time, place, and/or purpose, and purport to guarantee exclusivity in the use of public domain images.' The lawsuit also claims Getty has created 'a hostile environment for lawful users of public domain images' by allegedly sending them letters, via its License Compliance Services (LCS) subsidiary, accusing them of copyright infringement.

(edit: FWIW, this went to arbitration https://casetext.com/case/cixxfive-concepts-llc-v-getty-imag... and I can find nothing more on it since)

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#158
post #108

Earlier quoted context omitted.

The lawsuit doesn't even mention Google.

No, I did. What's your point?

I think it's a reference to Google's book scanning product, which is structurally similar: they use copyrighted material to provide a new kind of service, which contains an echo of the original material. The book scanning and the related search product is supposedly legal under U.S. copyright law.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#159

Earlier quoted context omitted.

Right that is the point of the parent comment - it’s not the book, it’s the amalgamation of all the discussions and content about the book. This case is probably dead in the water.

I'm not entirely up to speed on US law, but wouldn't OpenAI have to provide the court some kind of proof that they didn't use it in the training data during discovery?

Not a layer, but I believe the plaintiff (the author) would need to prove that it regurgitates their copyrighted work - otherwise it is possibly fair use. OpenAI does not need to prove anything, just defend their position at a reasonable level.

It’s not been decided if training a model on copyrighted works is “okay” or not as far as I know, but I expect it to be so, given that literally everyone does so at this point. It’s not like imagenet is copyright free, many of the images were/are.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#160

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Machine learning models have been trained with copyrighted data for a long time. Imagenet is full of copyrighted images, clearview literally just scanned the internet for faces, and I am sure there are other, older examples. I am unsure if this has been tested as fair use by a US court, but I am guessing it will be considered to be so if it is not already.
Post reply on HN