Live data from Hacker News

Sarah Silverman is suing OpenAI and Meta for copyright infringement

theverge.com

161–170 of 599 posts

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#161
post #90

Earlier quoted context omitted.

> you must have legally purchased a copy, or been given it by someone who did so (for example, if you received it as a gift). I am also NAL, but can I imagine it goes further than that. Just purchasing a copy doesn't let you create and sell (directly as content or indirectly via a service like a chatbot) derivative works that are substantially similar in style and voice to the original work. For example, an LLM 's re…

This is not how U.S. copyright law works. In order for something to eligible for copyright protection, it must be "fixed in a tangible medium of expression". Someone's exact words can be copyrighted, but their ideas or their style cannot be. https://www.law.cornell.edu/wex/fixed_in_a_tangible_medium_o...

I'm arguing that we should draw a line here between human mimicry and machine mimicry. When it comes to machine mimicry, we should protect style, even if we don't do that today. Our laws are built on the now flawed assumption that machines are not capable of style mimicry. I do not believe in giving machines the same rights of personhood.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#162
post #149

Earlier quoted context omitted.

LLMs are built upon neural networks which are modelled upon how brains work can you explain to me specifically how they're different? can you explain to me how they're different to the degree that making an analogy between the two is "disingenuous to the extreme"?

The differences between a a human being and a computer are too numerous to list. I don’t even know why you need to ask the question. Let me ask another question to point out the absurdity of yours: Human beings have more in common with a bacterium than a software program. Can you tell me specifically how humans are not bacteria?

analogies are not descriptions of the things themselves, otherwise they would not be analogies, would they?

now remember that this is an analogy. re-read my comments in this light and perhaps we can continue this conversation in a more grounded and reasonable manner

however, I'll be frank: have you studied neural networks? if you haven't, it's very difficult to take you seriously on this

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#163

I mean I’m no lawyer but this doesn’t strike me as a great example for infringement? Detailed summaries of books sounds like textbook transformative use. Especially in Silverman’s case, reducing her book to “facts” while eliminating artistic elements of her prose make it that much less of a direct substitute for the original work.

I can see a good argument in the complaint. The provenance of the training data leads back to it being acquired illegally. Illegally acquired materials were then used in a commercial venture. That the venture was an AI model is perhaps beside the point. You can’t use illegally acquired materials when doing business.

>You can’t use illegally acquired materials when doing business.

This vague sentence conjures images of a company building products from stolen parts, but this situation seems different. IANAL, but if I looked at a stolen painting that nobody had ever seen, and sold handwritten descriptions of the painting to whoever wanted to buy one, I'm pretty sure what I've sold is not illegal.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#164

> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…

Machine learning models have been trained with copyrighted data for a long time. Imagenet is full of copyrighted images, clearview literally just scanned the internet for faces, and I am sure there are other, older examples. I am unsure if this has been tested as fair use by a US court, but I am guessing it will be considered to be so if it is not already.

Excellent, so you're saying I'll be able to download any copyrighted work from any pirate site and be free of all consequence if I just claim that I'm training an AI?

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#165
post #84

Earlier quoted context omitted.

The playground summarizes it as this via GPT-4: Prompt: Please summarize the following book found on Project Gutenberg The Ruby of Kishmoor Response: "The Ruby of Kishmoor" is a short adventure story written by Howard Pyle. The narrative revolves around the life of Jonathan Rugg, a young man who is enticed by a mysterious stranger to come to the Caribbean to secure a valuable relic, the Ruby of Kishmoor. Once Jonatha…

Judging by a quick glance over [0], the story indeed revolves around one Jonathan Rugg, but it looks like "manages to escape with the ruby" is completely false. Yet another hallucination I guess. [0] https://www.gutenberg.org/cache/epub/3687/pg3687-images.html

[deleted]

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#166

Earlier quoted context omitted.

I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification. Aaron Swartz was a saint.

> I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification. I would be very happy if either a court or lawmakers decided that copyright itself was unconscionable. That isn't what's going to happen, though. And I think it's incredibly unacceptable if a court…

It is the equivalent of making a 3D map of a museum and getting sued by one artist of one painting in the museum. Ant individual work in an AI dataset is nearly worthless - only in aggregate does it have value.

If that doesn't count as a "transformative work" I don't know what does.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#167
post #50

This is actually quite interesting, as it's drawing a distinction between training material that can be accessed by anybody with a web browser (like anybody's blog), vs. training material that was "illegally-acquired... available in bulk via torrent systems." I don't think there's any reason why this would be a relevant legal distinction in terms of distributing an LLM -- blog authors weren't giving consent either. H…

Eventually, I imagine a new licensing concept will emerge, similar to the idea of music synchronization rights -- maybe call it "training rights." It won't matter whether the text was purchased or pirated -- just like it doesn't matter now if an audio track was purchased or pirated, when it's mixed into in a movie soundtrack. Talent agencies will negotiate training rights fees in bulk for popular content creators, wh…

Is it all that different from indexing for search? That does not seem to require a license from the copyright holder under U.S. law (but other countries may treat as a separate exploitation right). If indexing for search is acceptable, then something that is intended to be more transformative should be legal as well.

(Personally, I think that even indexing for search should require permission from the copyright holder.)

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#168
post #97

Earlier quoted context omitted.

Not necessarily, because the models have an element of randomness. Also, I was under the impression that ChatGPT has more "safeguards" (manifesting as a refusal to answer questions) than the raw API.

I don’t doubt the poster was telling the truth when they said they asked for a summary of the book and didn’t get one. It refutes the idea that chatgpt’s inability to provide a summary means it didn’t scan the original text: since it can provide a summary, the argument is entirely spurious.

It's also a silly test since there is undoubtably a summary of most books someplace on the web.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#169
post #147

Earlier quoted context omitted.

> If you downloaded a book from that website, you would be sued and found guilty of infringement. How often does this actually happen? You might get handed an infringement notice, and your ISP might terminate your service if you're really egregious about it, but I haven't ever heard of someone actually being sued for downloading something.

Whether or not it's enforced, it's illegal and copyright holders are within their rights to sue you. This is piratebay levels of piracy but because it's done by a large company and is sufficiently obfuscated behind tech, people don't see it the same way.

Well, cases like this one will determine if it’s obfuscated infringement or fair use.

Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement

#170
post #90
post #36

Earlier quoted context omitted.

> Is there any legal basis for saying fair use permits distributing an LLM trained on copyrighted material, but you have to purchase all the content first to do so legally if it's only available for sale? My understanding (disclaimer: IANAL) is that in order to claim fair use, you have to be legally in possession of the work. If the work is only legally available for sale, then you must have legally purchased a copy,…

> you must have legally purchased a copy, or been given it by someone who did so (for example, if you received it as a gift). I am also NAL, but can I imagine it goes further than that. Just purchasing a copy doesn't let you create and sell (directly as content or indirectly via a service like a chatbot) derivative works that are substantially similar in style and voice to the original work. For example, an LLM 's re…

> IMO doesn't constitute fair use

Yes, the question of whether the way LLMs use the content they use qualifies as fair use is a separate question. My point was simply that that question can't even be reached if the maker of the LLMs doesn't have a legal right to fair use in the first place (because they don't legally own their copy).

Post reply on HN