Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

151–160 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#151

I've been arguing since ChatGPT came out that LLMs should fall under fair use as a "transformative work". I'm not a lawyer and this is just my non-expert opinion, but it will be interesting to see what the legal system has to say about this.

What if I ask ChatGPT to print the article verbatim as sourced, from its own dataset?

It doesn't have database access to its own training dataset; it only has access to the weights it lossily-compressed that training dataset into.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#152

Earlier quoted context omitted.

in the same way that machines are not able to claim copyright, they aren't allowed to claim other legal rights either, like "fair use". The entity which owns ChatGPT is apparently maintaining a copy of the entirety of the New York Times archive within the ChatGPT knowledge base. That they extract some fair use snippets (they would claim) from it would still be fruit of a poisoned tree, no? (disclaimer: I'm pro AI, an…

I think there is some point between fifty years ago and last week in which the copyright for the content of newspapers should be public domain. That part of copyright needs to be fixed. Your creative work does deserve at least some period of exclusive rights for you. Definitely not so much that your grandchildren get to quibble about it well into retirement. But also whatever the number 3 or 4 most valuable company i…

> But also whatever the number 3 or 4 most valuable company in the world doesn’t get to scrape your content daily to repackage and sell as intelligent systems.

Here's a thing though: for 99%+ of that content, being turned into feedstock for ML model training is about the only valuable thing that came of its existence.

If it were not for world-ending danger of too smart an AI being developed too quickly, I'd vote for exempting ML training from copyright altogether, today - it's hard to overstate just how much more useful any copyrighted content is for society as LLM training data, than as whatever it was created for originally.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#153

in my head I like to think of web crawler search engines/search engine databases and LLMs as being somewhat similar. Search engines are ok if they just provide snippets with citations (urls), and they would be unacceptable if they provided large block quotes that removed the need to go to the original source to read the original expression of more complex ideas. A web-crawled LLM that lived within the same constraint…

I think it's different. LLMs can solve problems. Part of that problem-solving ability comes from training completely unrelated content such as NYT articles. GPT4 doesn't have to spit out NYT articles verbatim to have benefited from NYT articles. It uses NYT articles for every query.

Let's say I'm an academic; if my research, note-taking, and paper writing skills lead to fair-use, cited quotations where applicable, general knowledge not identified, and the creative aspects and unique conclusions creating the intriguing part of my work, that's copacetic. If I spit out (from memory, mind you) verbatim quotes and light rewordings of NY Times articles, that's not; "I don't remember where I got that material" doesn't cut it. My reading the NY Times every day for years because I judge it to be more literate and accurate than other sources, undoubtedly it has informed my thinking and style, but I don't need to acknowledge that.

If I use ChatGPT as a research tool, as long as it lives within the same parameters that I have to live within, I don't see a problem with its education/learning.

I understand that the NYTimes would like a slice of anything that comes out of the GPT but I'm talking about what seems reasonable. People who share their copyrighted material do not own all of the thinking that comes out of it; they own that expression of it, that is all.

Will AI destroy the economics of "writing" the way the web has killed newspapers? perhaps, perhaps we'll all benefit from and need a new model, but killing the new to keep the old on life support is not the way.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#154

The suit demonstrates instances where ChatGTP / Bing Copilot copy from the NYT verbatim. I think it is hard to argue that such copying constitutes "fair use". However, OAI/MS should be able to fix this within the current paradigm: Just learn to recognize and punish plagiarism via RLHF. However, the suit goes far beyond claiming that such copying violates their copyright: "Unauthorized copying of Times Works without p…

I think NYT is going to win. LLMs are arguably compressed data archives with weird algorithms. The fact that they will regularly regurgitate verbatim quotes of training data is evidence of this, as are the guardrails that try to prevent this. The second piece of evidence is this paper explained here https://www.hendrik-erz.de/post/why-gzip-just-beat-a-large-l... where instead of an LLM researchers used gzip compresse…

[deleted]

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#155
post #66

Earlier quoted context omitted.

Well yeah, copying a work and using it for its original expressive purpose isn’t fair use, no? You have to use it for a transformative purpose. Suppose I’m selling subscriptions to the New Jersey Times, a site which simply downloads New York Times articles and passes them through an autoencoder with some random noise. It serves the exact same purpose as the New York Times website, except I make the money. Is that fai…

If they could find a single person who in natural use (e.g. not as they were trying to gather data for this lawsuit) has ever actually used ChatGPT as a direct substitution for a NYT subscription, I'd support this lawsuit. But nobody would do that, because ChatGPT is a really shitty way to read NYT articles (it's stale, it can't reliably reproduce them, etc.). All that is valuable about it is the way that it transfor…

It’s more of a thought experiment. Here’s another with more commercial applications:

Suppose I start a service called “EastlawAI” by downloading the Westlaw database and hiring a team of comedians to write very funny lawyer jokes.

I take Westlaw cases and lawyer jokes and feed them to my autoencoder. I also learn a mapping from user queries to decoder inputs.

I sell an API and advertise it to startups as capable of answering any legal question in a funny way. Another company comes along with an API to make the output less funny.

Have I created a competitor to Westlaw by copying Westlaw’s works for their original expressive purpose and exposing it as an intermediary? Or have I simply trained the world’s most informative lawyer joke generator that some of my customers happen to use for legal analysis by layering other tools atop my output?

Did I need to download Westlaw cases to make my lawyer joke generator? Are the jokes a fair-use smokescreen for repackaging commercially valuable copyrighted data? Does my joke generator impact Westlaw in the market? Depends, right?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#156

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

> Typically I can't take a personal "tier" of a product and charge 3rd parties for derivatives of it. Say like VS Code. Can't you, though? I'd thought in general, it's a very important for the market to be able to do just that, otherwise everything gets gummed up in webs of exclusive contractual dependencies between established companies.

As I say, I don't really know. But then, this is exactly how SaaS licensing works. There may even be a free personal tier, where you can't sell products based on it, and a professional tier which may be very expensive indeed.

Typically providers of online databases go to some effort to stop people from sharing logins. Even from that point or view, I can imagine scraping articles and providing paraphrases of it for a fee is fishy.

All I'm saying, to some people it's obvious that the whole LLM on scraped Internet is fair use, to me it is not obvious.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#157
post #148

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

From what I can tell, this has nothing to do with LLMs at all. In the example in the article, the user is asking Bing to go fetch the contents of an article directly from the website, and print it out, which it dutifully does. Seems like the "problem" is that NYT etc gives privileged access to search engines for indexing their content, but then get upset when snippets of the indexed content is being shown to users wi…

I suppose that's a relatively easy thing to fix, technically. It proves, however, that th underlying LLM is trained on copyrighted data.

I'm not sure the problem goes away simply if the LLM in question (or any other one) gets some "no verbose regurgitation" filter.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#158

Earlier quoted context omitted.

I think it's different. LLMs can solve problems. Part of that problem-solving ability comes from training completely unrelated content such as NYT articles. GPT4 doesn't have to spit out NYT articles verbatim to have benefited from NYT articles. It uses NYT articles for every query.

Let's say I'm an academic; if my research, note-taking, and paper writing skills lead to fair-use, cited quotations where applicable, general knowledge not identified, and the creative aspects and unique conclusions creating the intriguing part of my work, that's copacetic. If I spit out (from memory, mind you) verbatim quotes and light rewordings of NY Times articles, that's not; "I don't remember where I got that m…

You're not replicating yourself millions of times and selling yourself for $20/month. If you are, then NYT might sue you too.

I'm not saying LLMs are by default, illegal. All I'm saying is that there is some merit to why NYT and content companies want a piece of the pie and think they deserve it.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#159

Earlier quoted context omitted.

Another factor to consider is that neural nets can function as lossy compression, which becomes extremely evident when using models that are overfit. Sometimes they're so overfit that the compression isn't even lossy, and the data is encoded verbatim in the NN.

Yes, but this then hits against learning/understanding and compression being fundamentally the same thing . I can't think of a better way to argue in favor of "it's fine if human does it, therefore it's fine if LLM does it", than from the "lossy compression" angle.

We can have different rules for humans than for machines. In fact, that happens all the time.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#160

Earlier quoted context omitted.

Maybe the bloom filter solution is enough, but I wonder. - Paraphrasing n=7 words (and quite a few more) within a sentence can easily be fair use. - As n gets big, the bloom filter has to also. If/when attribution is solved for LLMs (and not fake attribution like from Bing or Perplexity) then creators can be compensated when their works are used in AI outputs. If compensation is high enough this can greatly incentivi…

> if compensation is high enough Who pays the compensation? If it's the user, why wouldn't they just buy the authors work directly? Why go through the LLM middleman?

> If it's the user, why wouldn't they just buy the authors work directly? Why go through the LLM middleman?

If it's the user, why wouldn't they just buy the DVDs directly? Why go through the Netflix middleman?

A retort to this would be that both NYT and ChatGPT are on the internet, so it's no added fuss of hopping in my car, driving to Walmart, and picking up a DVD case. My response to it would be that both the LLM and Netflix are content aggregators to the user. I can read the NYT, or I can read the NYT summary on ChatGPT and ask it for life advice with my pet hamster, or ask it how to reverse a linked list in bash.

Post reply on HN