Earlier quoted context omitted.
Why do you say that? Commercial vs noncommercial use is a primary factor in the “purpose” prong of the fair use balancing test and a significant one in the “market effects” prong. That a use is noncommercial is often a deciding factor in the success of a fair use defense. GP is overstating it though, since it’s still one of many factors.
Your parent is more right than you. Weird Al has made a fantastic living copying music while only changing lyrics. He makes very heavy use of the satire plank of Fair Use. The “commercial” test is only part of the decision criteria for Fair Use.
NY Times copyright suit wants OpenAI to delete all GPT instances
791–800 of 921 posts
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#792Earlier quoted context omitted.
> if a work is purely derivative of a source work CliffNotes, Wikipedia, etc. have huge quantities of summarized copyrighted work.
First, you missed the "and". Do CliffNotes, Wikipedia, etc. substantially impact the market for the original work? For example CliffNotes does not - people who buy the CliffNotes version typically already have the original work as well (for example from coursework). And Wikipedia may well do more to interest people in the original work than to replace it. Second, you ignored the "purely derivative" bit. You have to l…
Is there data that supports this? I’d be interested to know what % of people who buy a Cliffs Notes have already _bought_ the original.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#793Earlier quoted context omitted.
> But I disagree with the underlying assumption that you can anthropomorphize LLMs. Gradient descent and backpropagation don't take place in the brain. LLMs "learn" in the same way that Excel sheets "learn". Backprop doesn't happen in us, but I think our neurones still do gradient descent – synapses that fire together, wire together. And ultimately, at the deepest level we can analyse, our brains' atoms are doing qua…
> We need to figure out what the deeper rules are that lead to the status quo, not merely mimic the superficial result. Sure, that's an interesting path of inquiry, and one should be free to understand themselves as being no different than a machine if they desire. But the objective of laws is the benefit of (at least some) humans, not machines covered in lab grown tissue. The process of being human is a big part of…
I think you're misapprehending — I mean an entity fully 3D printed out of tissue, no machinery (unless you're counting all biology as machinery, but I think you're not doing that).
I recon bio-printing is now where home computing was in the Apple 1 era, so this is a way off, but it's foreseeable.
> The process of being human is a big part of what makes us human.
Mmm. How much has that process that changed since the ancient world?
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#794Earlier quoted context omitted.
> Questions of fair use are famously gray, and anyone who declares something as "entirely fair use", with no caveats, is nearly always wrong except for the must obvious cases, which the given example is most definitely not. A judge has wide latitude in determining fair use. You're the one presenting unfounded claims with confidence here. There is well established case law about not being able to copyright facts. If y…
> You're the one presenting unfounded claims with confidence here. No, I'm not. On the contrary, I'm really looking forward to this case because I believe it will be a great test of a bunch of concepts that are totally novel in the world of copyright law as it applies to generative AI. The only things I am presenting with confidence are: 1. That anyone who declares that something is unambiguously fair use (or, contra…
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#795Earlier quoted context omitted.
Correct, just like it’s infringement to reproduce an article from memory using pen and paper intentionally. The person deciding to do that bears responsibility. OpenAI would be liable IFF they were intentionally facilitating that, instead of it being an undesired artifact from overfitting.
That's not true at all. Copyright infringement is a strict liability offense with no inquiry in to the state of the mind of the infringer from a liability perspective. The state of mind of the infringer is only relevant to the issue of willful infringement.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#796Earlier quoted context omitted.
> Is that fair use? As always, the answer is.. "it depends". I guess it depends mostly on the jurisdiction that applies to you. "Fair use" can have rather different legal meaning (or not exist at all) in different countries.
Fair use is specific to the US, as far as I'm aware. Moreover, Congress had to codify fair use (turn fair use common into statutory law in the form of 17 U.S. Code § 107) in order to make copyright statutes compatible with the First Amendment. Most other countries don't have freedom of expression and freedom of the press, so copyright law in a different country usually lacks a unifying exception test like fair use to…
This is demonstrably wrong. Many countries have both freedoms, albeit some have less strong protection than others.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#797If a NYT article says "Henry Kissenger was known to eat ice cream on a hot day" and our game outputs the same, it is purely by chance. It cannot be proven the output was copied verbatim from the NYT because the fragment "Henry Kissenger was known to" and "eat ice cream on a hot day" are not unique to the NYT or exclusive to it.
Is the NYT claiming ownership of the weights in LLMs?
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#798The suit demonstrates instances where ChatGTP / Bing Copilot copy from the NYT verbatim. I think it is hard to argue that such copying constitutes "fair use". However, OAI/MS should be able to fix this within the current paradigm: Just learn to recognize and punish plagiarism via RLHF. However, the suit goes far beyond claiming that such copying violates their copyright: "Unauthorized copying of Times Works without p…
> Just learn to recognize and punish plagiarism via RLHF. I'm not sure how your proposal would actually work. To recognize plagiarism during inference it needs to memorize harder. Kinda funny if it works though. We'd first train them to copy their training data verbatim , then train them not to. That is how it works, right? They're trained to copy their training data verbatim because that's the loss function. It's ju…
One thing you might do is use a full-text search database of the entire training data. If part of ChatGPT response is directly copied, give it the assignment of "please paraphrase this" and substitute the paraphrase into the response. This might slow ChatGPT down a lot - but it might not, I think an LLM is actually more computationally expensive than a full-text search by a lot.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#799Earlier quoted context omitted.
> What you described is entirely fair use, actually Just like during the pandemic how everyone became an epidemiologist, suddenly everyone's a copyright lawyer. I'll just dispute your assertion by saying: 1. Questions of fair use are famously gray, and anyone who declares something as "entirely fair use", with no caveats, is nearly always wrong except for the must obvious cases, which the given example is most defini…
> Questions of fair use are famously gray, and anyone who declares something as "entirely fair use", with no caveats, is nearly always wrong except for the must obvious cases, which the given example is most definitely not. A judge has wide latitude in determining fair use. You're the one presenting unfounded claims with confidence here. There is well established case law about not being able to copyright facts. If y…
Remember. The NY Times does not have a record of filing frivolous lawsuits. Particularly not against companies with deep pockets. So it is almost certainly true that a lawyer who knows the law better than you thinks that this has a real chance. So you should be looking for flaws in trivial defenses that you can think up, rather than assuming that you know best.
For example take your copyright facts defense. That would be great if the NY Times was a phone book. They aren't, in addition to facts they offer analysis, editorial positions, and so on. For example I just asked ChatGPT, "In 2016, did the New York Times generally support or oppose President Trump?" I got back an answer talking about various kinds of concerns that the New York Times had, including an editorial titled, "Why Donald Trump Should Not Be President". The copy that ChatGPT needed to have to do that has a lot more than just facts in it.
Now if you paraphrased the NY Times like ChatGPT did when it answered me, you'd have a perfect fair use defense. But you aren't doing it for money, you didn't make a copy of all the NY Times, you aren't destroying the market for the NY Times, and you're legally able to own copyright in your transformed work. OpenAI is doing it for money, did copy all of the NY Times, is seriously impacting the market for NY Times articles, and ChatGPT generated text does not get a copyright.
Fair use is filled with shades of grey. Even if ChatGPT appears to do the same thing that you do, it is far less clear that OpenAI will enjoy the same level of fair use defense.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#800Earlier quoted context omitted.
Why do you say that? Commercial vs noncommercial use is a primary factor in the “purpose” prong of the fair use balancing test and a significant one in the “market effects” prong. That a use is noncommercial is often a deciding factor in the success of a fair use defense. GP is overstating it though, since it’s still one of many factors.
Whether or not the use is commercial is certainly one of the considerations, but it's not the most significant one generally. There certainly can be specific cases where it's very significant, of course. But what I was arguing was that a use is not "fair use" merely because it's noncommercial in nature. I cannot make copies of movies and give them away on the street for free and successfully claim "fair use".