Earlier quoted context omitted.
And you think that it would be impossible to train a model to avoid outputs that are substantially similar to training data?
I certainly don't think it's impossible, but I think it is hard problem that won't be solved in the immediate future, and creators of data used for training are right to seek to stop wide availability of LLMs that regurgitate information they worked hard to obtain.
NY Times copyright suit wants OpenAI to delete all GPT instances
821–830 of 921 posts
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#822Earlier quoted context omitted.
That's not true at all. Copyright infringement is a strict liability offense with no inquiry in to the state of the mind of the infringer from a liability perspective. The state of mind of the infringer is only relevant to the issue of willful infringement.
Willful vs less-than-willful infringement are definitely two separate types of offences, as indicated by the difference of penalty.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#823Earlier quoted context omitted.
Because anyone that is familiar with fair use knows that the purpose prong and the commerciality aspect of it is not one of the more important prongs of the fair use analysis, whereas transformation is. Transformation adjusts what is a purpose that falls under fair use. Did you read Warhol??
Yes. Warhol is an example where the commercial nature of the secondary use was the deciding factor in its failure to pass the purpose prong. > In sum, if an original work and secondary use share the same or highly similar purposes, and the secondary use is commercial, the first fair use factor is likely to weigh against fair use, absent some other justification for copying. (P4). It’s very likely that a noncommercial…
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#824Earlier quoted context omitted.
Legality aside I think copyright of digital things in the digital age is a net negative to humanity.
Completely agree. Copyright should be abolished. All intellectual work is information, information is just bits and bits are just numbers. It's quite simply delusional to believe you can own numbers in the 21st century, the age of information and ubiquitous globally networked pocket supercomputers. This is just a felony contempt of business model issue. Computers invalidated their business models and they're doing ev…
What do you propose is the business model for artists in the absence of copyright?
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#825Earlier quoted context omitted.
Completely agree. Copyright should be abolished. All intellectual work is information, information is just bits and bits are just numbers. It's quite simply delusional to believe you can own numbers in the 21st century, the age of information and ubiquitous globally networked pocket supercomputers. This is just a felony contempt of business model issue. Computers invalidated their business models and they're doing ev…
That seems like a great way to destroy what is left of art as we know it. Anna Karenina is just numbers. In The Mood For Love is just numbers. Right. What do you propose is the business model for artists in the absence of copyright?
We must strengthen these business models that don't depend on artificial scarcity because this number selling nonsense was over the second computers were invented. It's as dumb as asserting that you need permission to use memcpy or the mov CPU instruction.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#826If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#827If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#828Earlier quoted context omitted.
I think the issue is that they trained ChatGPT on the New York Times' proprietary IP without paying licensing fees and, the Times argues, that is illegal. By way of proof the Times has examples of ChatGPT dumping out articles verbatim.
IMO it's pretty hard to training an LLM isn't a transformative use. It's clearly not just copying, or even excerpting. Even if it was just compression (and it's not), they're only providing model output not distributing the "compressed" NYT articles. Yielding verbatim snippets of copyrighted content is a problem for OpenAI though.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#829Earlier quoted context omitted.
I don't think you could use RLHF to stop plagerism. RLHF can be used to teach what "angry response" is because you look at the text itself for qualities. A plagerized text doesn't have any special qualities aside from "existing already", which you can only determine by looking at the world. One thing you might do is use a full-text search database of the entire training data. If part of ChatGPT response is directly c…
I agree that this sketch comes closer to working in practice than simple RLHF. In my earlier comment I was imagining bringing in some auxiliary data like you describe to detect plagarism and then using RL to teach the model not to do it.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#830Earlier quoted context omitted.
Why do you say that? Commercial vs noncommercial use is a primary factor in the “purpose” prong of the fair use balancing test and a significant one in the “market effects” prong. That a use is noncommercial is often a deciding factor in the success of a fair use defense. GP is overstating it though, since it’s still one of many factors.
Your parent is more right than you. Weird Al has made a fantastic living copying music while only changing lyrics. He makes very heavy use of the satire plank of Fair Use. The “commercial” test is only part of the decision criteria for Fair Use.
Not only does he get permission from the original authors, he also pays royalties to the authors despite legally not having to do so.