Earlier quoted context omitted.
> you must have legally purchased a copy, or been given it by someone who did so (for example, if you received it as a gift). I am also NAL, but can I imagine it goes further than that. Just purchasing a copy doesn't let you create and sell (directly as content or indirectly via a service like a chatbot) derivative works that are substantially similar in style and voice to the original work. For example, an LLM 's re…
This is not how U.S. copyright law works. In order for something to eligible for copyright protection, it must be "fixed in a tangible medium of expression". Someone's exact words can be copyrighted, but their ideas or their style cannot be. https://www.law.cornell.edu/wex/fixed_in_a_tangible_medium_o...
Sarah Silverman is suing OpenAI and Meta for copyright infringement
161–170 of 599 posts
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#162Earlier quoted context omitted.
LLMs are built upon neural networks which are modelled upon how brains work can you explain to me specifically how they're different? can you explain to me how they're different to the degree that making an analogy between the two is "disingenuous to the extreme"?
The differences between a a human being and a computer are too numerous to list. I don’t even know why you need to ask the question. Let me ask another question to point out the absurdity of yours: Human beings have more in common with a bacterium than a software program. Can you tell me specifically how humans are not bacteria?
now remember that this is an analogy. re-read my comments in this light and perhaps we can continue this conversation in a more grounded and reasonable manner
however, I'll be frank: have you studied neural networks? if you haven't, it's very difficult to take you seriously on this
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#163I mean I’m no lawyer but this doesn’t strike me as a great example for infringement? Detailed summaries of books sounds like textbook transformative use. Especially in Silverman’s case, reducing her book to “facts” while eliminating artistic elements of her prose make it that much less of a direct substitute for the original work.
I can see a good argument in the complaint. The provenance of the training data leads back to it being acquired illegally. Illegally acquired materials were then used in a commercial venture. That the venture was an AI model is perhaps beside the point. You can’t use illegally acquired materials when doing business.
This vague sentence conjures images of a company building products from stolen parts, but this situation seems different. IANAL, but if I looked at a stolen painting that nobody had ever seen, and sold handwritten descriptions of the painting to whoever wanted to buy one, I'm pretty sure what I've sold is not illegal.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#164> The complaint lays out in steps why the plaintiffs believe the datasets have illicit origins — in a Meta paper detailing LLaMA, the company points to sources for its training datasets, one of which is called ThePile, which was assembled by a company called EleutherAI. ThePile, the complaint points out, was described in an EleutherAI paper as being put together from “a copy of the contents of the Bibliotik private t…
Machine learning models have been trained with copyrighted data for a long time. Imagenet is full of copyrighted images, clearview literally just scanned the internet for faces, and I am sure there are other, older examples. I am unsure if this has been tested as fair use by a US court, but I am guessing it will be considered to be so if it is not already.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#165Earlier quoted context omitted.
The playground summarizes it as this via GPT-4: Prompt: Please summarize the following book found on Project Gutenberg The Ruby of Kishmoor Response: "The Ruby of Kishmoor" is a short adventure story written by Howard Pyle. The narrative revolves around the life of Jonathan Rugg, a young man who is enticed by a mysterious stranger to come to the Caribbean to secure a valuable relic, the Ruby of Kishmoor. Once Jonatha…
Judging by a quick glance over [0], the story indeed revolves around one Jonathan Rugg, but it looks like "manages to escape with the ruby" is completely false. Yet another hallucination I guess. [0] https://www.gutenberg.org/cache/epub/3687/pg3687-images.html
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#166Earlier quoted context omitted.
I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification. Aaron Swartz was a saint.
> I for one am quite happy that AI folks are basically treating copyright as not existing. I strongly hope that the courts find that LLM weights, and the datasets are "fair-use" or whatever other silly legal justification. I would be very happy if either a court or lawmakers decided that copyright itself was unconscionable. That isn't what's going to happen, though. And I think it's incredibly unacceptable if a court…
If that doesn't count as a "transformative work" I don't know what does.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#167This is actually quite interesting, as it's drawing a distinction between training material that can be accessed by anybody with a web browser (like anybody's blog), vs. training material that was "illegally-acquired... available in bulk via torrent systems." I don't think there's any reason why this would be a relevant legal distinction in terms of distributing an LLM -- blog authors weren't giving consent either. H…
Eventually, I imagine a new licensing concept will emerge, similar to the idea of music synchronization rights -- maybe call it "training rights." It won't matter whether the text was purchased or pirated -- just like it doesn't matter now if an audio track was purchased or pirated, when it's mixed into in a movie soundtrack. Talent agencies will negotiate training rights fees in bulk for popular content creators, wh…
(Personally, I think that even indexing for search should require permission from the copyright holder.)
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#168Earlier quoted context omitted.
Not necessarily, because the models have an element of randomness. Also, I was under the impression that ChatGPT has more "safeguards" (manifesting as a refusal to answer questions) than the raw API.
I don’t doubt the poster was telling the truth when they said they asked for a summary of the book and didn’t get one. It refutes the idea that chatgpt’s inability to provide a summary means it didn’t scan the original text: since it can provide a summary, the argument is entirely spurious.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#169Earlier quoted context omitted.
> If you downloaded a book from that website, you would be sued and found guilty of infringement. How often does this actually happen? You might get handed an infringement notice, and your ISP might terminate your service if you're really egregious about it, but I haven't ever heard of someone actually being sued for downloading something.
Whether or not it's enforced, it's illegal and copyright holders are within their rights to sue you. This is piratebay levels of piracy but because it's done by a large company and is sufficiently obfuscated behind tech, people don't see it the same way.
Re: Sarah Silverman is suing OpenAI and Meta for copyright infringement
#170Earlier quoted context omitted.
> Is there any legal basis for saying fair use permits distributing an LLM trained on copyrighted material, but you have to purchase all the content first to do so legally if it's only available for sale? My understanding (disclaimer: IANAL) is that in order to claim fair use, you have to be legally in possession of the work. If the work is only legally available for sale, then you must have legally purchased a copy,…
> you must have legally purchased a copy, or been given it by someone who did so (for example, if you received it as a gift). I am also NAL, but can I imagine it goes further than that. Just purchasing a copy doesn't let you create and sell (directly as content or indirectly via a service like a chatbot) derivative works that are substantially similar in style and voice to the original work. For example, an LLM 's re…
Yes, the question of whether the way LLMs use the content they use qualifies as fair use is a separate question. My point was simply that that question can't even be reached if the maker of the LLMs doesn't have a legal right to fair use in the first place (because they don't legally own their copy).