Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

811–820 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#811

Earlier quoted context omitted.

Legality aside I think copyright of digital things in the digital age is a net negative to humanity.

Completely agree. Copyright should be abolished. All intellectual work is information, information is just bits and bits are just numbers. It's quite simply delusional to believe you can own numbers in the 21st century, the age of information and ubiquitous globally networked pocket supercomputers. This is just a felony contempt of business model issue. Computers invalidated their business models and they're doing ev…

This goes too far. Digital media are not only long series of numbers. They are often difficult-to-create expressions in image, video, and even interactive forms; regardless of their serialization format.

Books are just strings of letters, yet copyright has still been useful to increase the volume and utility of books.

All that said, I do find the life+70y an absurdly long time.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#812
post #737

Related. Others? NYT sues OpenAI, Microsoft over 'millions of articles' used to train ChatGPT - https://news.ycombinator.com/item?id=38784194 - Dec 2023 (80 comments) The New York Times is suing OpenAI and Microsoft for copyright infringement - https://news.ycombinator.com/item?id=38781941 - Dec 2023 (837 comments) The Times Sues OpenAI and Microsoft Over A.I.’s Use of Copyrighted Work - https://news.ycombinator.com/…

New York Times Sues Microsoft and OpenAI for 'Billions' - https://news.ycombinator.com/item?id=38791368 - Dec 2023 (1 comment) New York Times Sues Microsoft and OpenAI, Alleging Copyright Infringement - https://news.ycombinator.com/item?id=38781718 - Dec 2023 (1 comment) New York Times sues Microsoft and OpenAI over copyright infringement - https://news.ycombinator.com/item?id=38781908 - Dec 2023 (2 comments) New Yor…

I already went through those links and merged the comments into the above threads, so there's nothing in any of those (besides archive links to the articles).

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#813

The NYT is preparing for a tsunami by building a sandcastle. Big picture, this suit won’t matter, for so many reasons. To enumerate a few: 1. Next gen LLMs will be trained exclusively on “synthetic”/public data. GPT-4V can easily whitewash its entire copyrighted training corpus to be unrecognizably distinct (say reworded by 40%, authors/sources stripped, etc). Ergo there will be no copyright material for GPT-5 to reg…

I think it can be simultaneously true that NYT is accurate in their complaint, while having no legal remedy for this and that there shouldn’t be.

There are plenty of large companies in other sectors that acknowledge there are limited legal remedies for them if someone copies some aspect of their business or name.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#814
post #810

Earlier quoted context omitted.

Copyright law doesn't mention opt outs or search engine snippet controls. It's not clear to me that robots.txt is the singular thing that makes Google legal. In US copyright law facts cannot be copyrighted, so copyright on factual content like newspaper articles is limited. Simply replacing a few words wouldn't work, but I am certain that GPT-4 is capable of paraphrasing factual content at a level that would not be c…

That’s not the only reason. Google search is also transformative and non competitive with the underlying publications. And that is why the opt out is important. If you feel google competes with your site you don’t have to sue Google: just tell them to to away

Transformative yes, so is ChatGPT. Much more so actually. Non-competitive is debatable. Especially with the instant answers Google has in addition to regular snippets which can also obviate the need to visit a site. I have a hard time seeing ChatGPT as competing with newspapers more than Google Search does.

Nobody is seriously going to ChatGPT and trying to trick it into regurgitating old NYT articles as an alternative to paying for access to NYT's archives. Meanwhile, newspapers went as far as getting the laws changed in several countries because they felt Google was competing with them too much and didn't like the fact that it was legal.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#815
post #793
post #720

Earlier quoted context omitted.

> We need to figure out what the deeper rules are that lead to the status quo, not merely mimic the superficial result. Sure, that's an interesting path of inquiry, and one should be free to understand themselves as being no different than a machine if they desire. But the objective of laws is the benefit of (at least some) humans, not machines covered in lab grown tissue. The process of being human is a big part of…

> machines covered in lab grown tissue I think you're misapprehending — I mean an entity fully 3D printed out of tissue, no machinery (unless you're counting all biology as machinery, but I think you're not doing that). I recon bio-printing is now where home computing was in the Apple 1 era, so this is a way off, but it's foreseeable . > The process of being human is a big part of what makes us human. Mmm. How much h…

I recon bio-printing is now where home computing was in the Apple 1 era

How do you recon that, Apple 1 was Turing complete. We haven't printed life, that would be a tremendous accomplishment.

I think we're closer to Edison inventing a lightbulb as a step to computers being possible. Printing a conscious thing, at all, would be like the transistor. An Apple 1 analogue wouldn't be likely because of the terrible ethics of a "shitty" printed human.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#816

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

> To keep things simple, let's say I never regurgitate chunks of verbatim NYT articles, maybe quite short snippets. You just described Google. When you think about it, it's surprising that Google is legal. However, it is well established that what Google does is perfectly legal. Remember that internally Google keeps and uses complete verbatim copies of every web page they index. Yes, Google offers a link to the sourc…

The reason why Google keeping entire digital copies of other people's copyrighted works is legal is because copyright is all about distribution rights. Any person can possess the entire works of Disney (without paying for them), for example and as long as they do not distribute those works they're 100% in the clear.

Possession is not a crime when it comes to copyright. It's not like physical things (e.g. drugs or guns) at all. This is why comparing copyright violations to theft is silly.

ChatGPT can absolutely keep verbatim copies of the entire works of basically anything and not run afoul of the law. When it regurgitates a small part of an article that's covered by fair use in theory but the truth is that fair use can only be determined by a judge in a court of law when someone is sued. It cannot be determined with any sort of certainty ahead of time. It's a legal defense, nothing more.

Summarizing content has been legal forever as well (see the other posts here talking about Cliff Notes and some similar products). That's not even fair use that's just like, people's opinions, man (legally speaking).

I don't think the NYT will get what they want out of this at all.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#817

Earlier quoted context omitted.

> Just learn to recognize and punish plagiarism via RLHF. I'm not sure how your proposal would actually work. To recognize plagiarism during inference it needs to memorize harder. Kinda funny if it works though. We'd first train them to copy their training data verbatim , then train them not to. That is how it works, right? They're trained to copy their training data verbatim because that's the loss function. It's ju…

I don't think you could use RLHF to stop plagerism. RLHF can be used to teach what "angry response" is because you look at the text itself for qualities. A plagerized text doesn't have any special qualities aside from "existing already", which you can only determine by looking at the world. One thing you might do is use a full-text search database of the entire training data. If part of ChatGPT response is directly c…

I agree that this sketch comes closer to working in practice than simple RLHF. In my earlier comment I was imagining bringing in some auxiliary data like you describe to detect plagarism and then using RL to teach the model not to do it.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#818
post #762

Earlier quoted context omitted.

> What you described is entirely fair use, actually Just like during the pandemic how everyone became an epidemiologist, suddenly everyone's a copyright lawyer. I'll just dispute your assertion by saying: 1. Questions of fair use are famously gray, and anyone who declares something as "entirely fair use", with no caveats, is nearly always wrong except for the must obvious cases, which the given example is most defini…

>if a work is purely derivative of a source work This is the weakest part of the case(s) against OpenAI. "Derivative work" is a legal term of art meaning a direct adaptation, like writing a screenplay of a book or translating a book into another language. NYT has a stronger case than Sarah Silverman here because they can show actual 'memorized' text rather than just summarization, but given that those memorizations a…

A question is whether the new model still intrinsically embeds the source text, but this is later filtered in the output, or if it no longer embeds the text at all.

The latter is more defensible.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#819
post #190

Earlier quoted context omitted.

Isn't it totally normal to write articles / blog posts that effectively summarize, and often quote from, news articles?

I think the issue is that they trained ChatGPT on the New York Times' proprietary IP without paying licensing fees and, the Times argues, that is illegal. By way of proof the Times has examples of ChatGPT dumping out articles verbatim.

IMO it's pretty hard to training an LLM isn't a transformative use. It's clearly not just copying, or even excerpting. Even if it was just compression (and it's not), they're only providing model output not distributing the "compressed" NYT articles.

Yielding verbatim snippets of copyrighted content is a problem for OpenAI though.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#820
post #772

Earlier quoted context omitted.

> if a work is purely derivative of a source work CliffNotes, Wikipedia, etc. have huge quantities of summarized copyrighted work.

First, you missed the "and". Do CliffNotes, Wikipedia, etc. substantially impact the market for the original work? For example CliffNotes does not - people who buy the CliffNotes version typically already have the original work as well (for example from coursework). And Wikipedia may well do more to interest people in the original work than to replace it. Second, you ignored the "purely derivative" bit. You have to l…

> Do CliffNotes, Wikipedia, etc. substantially impact the market for the original work?

Yes.

For example, Wikipedia cites many research journals that otherwise are available only by subscription.

Prior to Wikipedia, gated information centers were the norm.

Post reply on HN