Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

621–630 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#621
post #514

Earlier quoted context omitted.

> why we feel it's OK to pirate news articles, but not other IP Who thinks this? I don't. I think copyright is wrong across the board. I would love if the same pattern of posting archive'd articles held for books, movies, et cetera. I would love to change my mind on this, as it is a very unpopular opinion to have. But I have _never_ seen a morally or scientifically sound argument in favor of copyright law, and I've s…

You can self-publish. Oh, you want to be able to publish other people’s work, and without their permission? How does that benefit the author?

> How does that benefit the author?

You speak of "the author". But the current system does not benefit "the author". 1% of authors profit off copyright. 99% lose money on copyright (they pay more for copyrighted media than they earn from it).

Your question should be "How does that benefit monopolist authors"?

I agree, my idea would not benefit monopolist authors. They would lose the bulk of their revenue stream.

But it would benefit the average author whose cost of living would fall and information would start serving them more than serving business.

I am not downplaying the talent and hard work of successful monopolist authors. But I do not think the works they create are worth everyone giving up their rights to reshare and remix information. I believe the world would look very different post-IP. You'd probably have a new profession--small independent librarians (similar to data hoarders today)--who would help their local communities maximize the value they got from humanity's best information.

Maybe I'm wrong! Maybe the information ecosystem is better controlled and the genetic differences of monopolist authors are so stark that without the subsidies to this gifted class we'd all be worse off. But that's an argument based on outcomes and not principles.

> without their permission

The oxygen I'm breathing right now mostly was created by trees on land owned by others. But I don't ask for their permission to breath. Some things are just not natural.

I am not saying plagiarize. It is always the right thing to do to link back and/or credit the source. But needing to ask permission to republish something seems to go against natural laws.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#622

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

> To keep things simple, let's say I never regurgitate chunks of verbatim NYT articles, maybe quite short snippets. You just described Google. When you think about it, it's surprising that Google is legal. However, it is well established that what Google does is perfectly legal. Remember that internally Google keeps and uses complete verbatim copies of every web page they index. Yes, Google offers a link to the sourc…

> However, it is well established that what Google does is perfectly legal.

Google has a wide range of products and shakedowns. Not all of them are "perfectly" legal: Google is being challenged in court over some of their shakedowns and products practices.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#623

The suit demonstrates instances where ChatGTP / Bing Copilot copy from the NYT verbatim. I think it is hard to argue that such copying constitutes "fair use". However, OAI/MS should be able to fix this within the current paradigm: Just learn to recognize and punish plagiarism via RLHF. However, the suit goes far beyond claiming that such copying violates their copyright: "Unauthorized copying of Times Works without p…

Adding an extra constraint of no copying verbatim from a very large and relevant corpus will be hard to guarantee without enormous databases of copyrighted content (which might not be legal to hold) and add an extra objective to a system with many often contradictory goals. I don’t think that’s the technology-sound solution or one in the interest of anyone involved. It’s much more relevant to license content from as many newspapers as possible, recognize when references are relevant, and quote them either explicitly verbatim if that’s the best answer or adapt (translate, simplify, add context) when appropriate.

I feel like the NYTimes is asking for deletion as a negotiation tactic to force OpenAI to give them enough money to pay for their journalism (I am not sure who would subscribe to NYTimes if you can get as much through OpenAI, but I am open to registering extra to pay for their work).

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#624

Earlier quoted context omitted.

> To keep things simple, let's say I never regurgitate chunks of verbatim NYT articles, maybe quite short snippets. You just described Google. When you think about it, it's surprising that Google is legal. However, it is well established that what Google does is perfectly legal. Remember that internally Google keeps and uses complete verbatim copies of every web page they index. Yes, Google offers a link to the sourc…

> However, it is well established that what Google does is perfectly legal. Google has a wide range of products and shakedowns. Not all of them are "perfectly" legal: Google is being challenged in court over some of their shakedowns and products practices.

I am clearly talking about the web search engine in the context of copyright. Other products or legal concerns like antitrust are completely irrelevant here.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#625

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

> To keep things simple, let's say I never regurgitate chunks of verbatim NYT articles, maybe quite short snippets. You just described Google. When you think about it, it's surprising that Google is legal. However, it is well established that what Google does is perfectly legal. Remember that internally Google keeps and uses complete verbatim copies of every web page they index. Yes, Google offers a link to the sourc…

Any publisher can opt out of google. Publisher also have substantial control over titles and snippets shown in google, whether an article appears in google news, etc

Paraphrasing is also known as cloning and is often a copyright violation

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#626
If they don't let AIs to be trained on a maximum of data as possible, those AIs will be less "good" than the ones trained without constraints like you will have in China or elsewhere, and people will mechanically start using the later.

Unless they engage in massive IP and DNS banning, geolocation based, that forced upon all internet users and "external" users.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#627

The suit demonstrates instances where ChatGTP / Bing Copilot copy from the NYT verbatim. I think it is hard to argue that such copying constitutes "fair use". However, OAI/MS should be able to fix this within the current paradigm: Just learn to recognize and punish plagiarism via RLHF. However, the suit goes far beyond claiming that such copying violates their copyright: "Unauthorized copying of Times Works without p…

> This is a strong claim that just downloading articles into training data is what violates the copyright. That GTP outputs verbatim copies is a red herring.

It's the other way around. There is no infringement if the model output is not substantially similar to a work in the training set [1]:

> To win a claim of copyright infringement in civil or criminal court, a plaintiff must show he or she owns a valid copyright, the defendant actually copied the work, and the level of copying amounts to misappropriation.

The questions are, which parties should bear liability when the model creates infringing outputs, and how should that liability be split among the parties? Given that getting an infringing output likely requires the prompt to reference an existing work (which is what's happening in the article), an author of a work, an element in an existing work, or a characteristic/style strongly associated with certain works/authors, I believe that the user who makes the prompt should bear most of the liability should the user choose to publish an infringing output in a way that doesn't fall under fair use. (AI companies should not be publishing model outputs by default.)

[1] https://en.wikipedia.org/wiki/Substantial_similarity#Substan...

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#628

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

I agree with your IANAL take, but what about a situation with an extra level of indirection? So the service never reads actual NYT articles, but only reads blog/forum posts about NYT articles, and derives what is in the article from conversations about the article by people who have read it. Is that legal now?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#629

Earlier quoted context omitted.

Is there some LLM meta where understanding and compression are argued to be the same thing I’m not aware of? Anyone got more details on this? Superficially it sounds like total BS; a highly compressed zip file does not exhibit any characteristics of learning. Algorithmically derived highly compressed video streams do not exhibit characteristics of learning. ? I’ve vaguely heard the learning can be considered to exhib…

Suppose you wanted to train an LLM to do addition. An LLM has limited parameters. If an LLM had infinite parameters it could just memorize the results of every single addition question in existence and could not claim to have understood anything. Because it has finite parameters, if an LLM wants to get a lower loss on all addition questions, it needs to come up with a general algorithm to perform addition. Indeed, Ne…

It’s exactly this kind of thinking that underlies lossless text compression (not exactly what a transformer guarantees but often what happens). For that reason, some people thought it would be fun to combine zip and transformers. https://openreview.net/forum?id=hO0c2tG2xL

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#630
post #615

Isn't the fundamental issue here that the NYT was available in Common Crawl? If they didn't want to share their content, why did they allow it to be scraped? If they did want to share their content, why do they care (hint: $88 billion)? Or is it that they wanted to share their content with Google and other search engines in order to bring in readers but now that an AI was trained on it they are angry? What wrong thin…

If you read the complaint, it explains this pretty well. The use of copyrighted content by search engines is fundamentally different from the way LLMs use that same content. The former directs traffic (and therefore $$) to the publisher, the latter keeps the traffic for itself.

The legal misconception I want to flag in your logic is the notion that all uses of the Common Crawl are equally infringing/non-infringing. If you use the Common Crawl to create a list of how often every word in English appears on the internet, that’s unquestionably transformative use. But if you use it to host a mirror of the NYT website with free articles, that’s definitely infringement. The legality of scraping is one matter, and the legality of what you do with the scraped content is quite another.

Post reply on HN