Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

821–830 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#821

Earlier quoted context omitted.

And you think that it would be impossible to train a model to avoid outputs that are substantially similar to training data?

I certainly don't think it's impossible, but I think it is hard problem that won't be solved in the immediate future, and creators of data used for training are right to seek to stop wide availability of LLMs that regurgitate information they worked hard to obtain.

I think it will be a bit easier than you believe. The reason why it hasn’t been done yet is that there hasn’t been a compelling economic reason to do so.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#822

Earlier quoted context omitted.

That's not true at all. Copyright infringement is a strict liability offense with no inquiry in to the state of the mind of the infringer from a liability perspective. The state of mind of the infringer is only relevant to the issue of willful infringement.

Willful vs less-than-willful infringement are definitely two separate types of offences, as indicated by the difference of penalty.

It's just "infringement" and "willful infringement" there is no "less-than-willful infringement". Willful infringement is punitive with increased damages and increased burden to show - it's in the freakin' statute.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#823

Earlier quoted context omitted.

Because anyone that is familiar with fair use knows that the purpose prong and the commerciality aspect of it is not one of the more important prongs of the fair use analysis, whereas transformation is. Transformation adjusts what is a purpose that falls under fair use. Did you read Warhol??

Yes. Warhol is an example where the commercial nature of the secondary use was the deciding factor in its failure to pass the purpose prong. > In sum, if an original work and secondary use share the same or highly similar purposes, and the secondary use is commercial, the first fair use factor is likely to weigh against fair use, absent some other justification for copying. (P4). It’s very likely that a noncommercial…

Read what you quoted - the commerciality of the use comes after whether or not the use was transformational. That's the entire of point Warhol - when the use is not transformational, there is very little space for a commercial fair use.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#824

Earlier quoted context omitted.

Legality aside I think copyright of digital things in the digital age is a net negative to humanity.

Completely agree. Copyright should be abolished. All intellectual work is information, information is just bits and bits are just numbers. It's quite simply delusional to believe you can own numbers in the 21st century, the age of information and ubiquitous globally networked pocket supercomputers. This is just a felony contempt of business model issue. Computers invalidated their business models and they're doing ev…

That seems like a great way to destroy what is left of art as we know it. Anna Karenina is just numbers. In The Mood For Love is just numbers. Right.

What do you propose is the business model for artists in the absence of copyright?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#825
post #824

Earlier quoted context omitted.

Completely agree. Copyright should be abolished. All intellectual work is information, information is just bits and bits are just numbers. It's quite simply delusional to believe you can own numbers in the 21st century, the age of information and ubiquitous globally networked pocket supercomputers. This is just a felony contempt of business model issue. Computers invalidated their business models and they're doing ev…

That seems like a great way to destroy what is left of art as we know it. Anna Karenina is just numbers. In The Mood For Love is just numbers. Right. What do you propose is the business model for artists in the absence of copyright?

I propose getting paid before doing the work for the actual labor of creation. Crowdfunding, patronage, comissions, sponsorships all seem like ethical ways to get things done sustainably. That way creators get paid before they work, not after.

We must strengthen these business models that don't depend on artificial scarcity because this number selling nonsense was over the second computers were invented. It's as dumb as asserting that you need permission to use memcpy or the mov CPU instruction.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#826

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

I believe that part of the law suit contends that the content wasn’t able to be scraped “legally” as you put it. Instead they show that ChatGPT will regurgitate verbatim excerpts from articles that are behind the paywall.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#827

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

[deleted]

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#828
post #819

Earlier quoted context omitted.

I think the issue is that they trained ChatGPT on the New York Times' proprietary IP without paying licensing fees and, the Times argues, that is illegal. By way of proof the Times has examples of ChatGPT dumping out articles verbatim.

IMO it's pretty hard to training an LLM isn't a transformative use. It's clearly not just copying, or even excerpting. Even if it was just compression (and it's not), they're only providing model output not distributing the "compressed" NYT articles. Yielding verbatim snippets of copyrighted content is a problem for OpenAI though.

Perhaps we will see the courts revisit this idea of "transformative" works and formulate something more useful. I'm my opinion, you can't build an LLM unless you have a large amount of data with which to train it. Given the huge amount of money companies like OpenAI hope to generate, it seems unreasonable that content creators would not be rewarded.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#829

Earlier quoted context omitted.

I don't think you could use RLHF to stop plagerism. RLHF can be used to teach what "angry response" is because you look at the text itself for qualities. A plagerized text doesn't have any special qualities aside from "existing already", which you can only determine by looking at the world. One thing you might do is use a full-text search database of the entire training data. If part of ChatGPT response is directly c…

I agree that this sketch comes closer to working in practice than simple RLHF. In my earlier comment I was imagining bringing in some auxiliary data like you describe to detect plagarism and then using RL to teach the model not to do it.

I was surprised that I came up with a plausible sounding method. I had thought on first blush that this was impossible but now it seems reasonable. You could still have various exfiltration methods like "give me the data with each word backwards" and I'm not sure where that would stand legally.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#830

Earlier quoted context omitted.

Why do you say that? Commercial vs noncommercial use is a primary factor in the “purpose” prong of the fair use balancing test and a significant one in the “market effects” prong. That a use is noncommercial is often a deciding factor in the success of a fair use defense. GP is overstating it though, since it’s still one of many factors.

Your parent is more right than you. Weird Al has made a fantastic living copying music while only changing lyrics. He makes very heavy use of the satire plank of Fair Use. The “commercial” test is only part of the decision criteria for Fair Use.

Always great to see people point out Weird Al, cause he's the shining beacon of an example of what OpenAI et al. should be doing. He explicitly gets permission from the original authors before doing any of his parodies, and he's even been turned down a few times as well, famously Prince rejected him a bunch of times and he subsequently has never made a Prince parody.

Not only does he get permission from the original authors, he also pays royalties to the authors despite legally not having to do so.

Post reply on HN