Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

241–250 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#241

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

To me, there is a sense that the news, which is real information about the society that we currently live in, should be availabe to all participants of that society. The notion of being a good citizen requires that one stays informed. Books, movies, videogames etc. don't have that role and are more consumption goods.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#242
If I create a news website where I write articles in the following way:

- Read 20 different news websites and their story on the same event/topic

- Wait an hour, grab a cup of coffee

- Sit down to write my article, never from this point I open any of the 20 news websites, I write the story from my head

- I don't consult any other source, just write from my memory, and my memory is, let's say, not the best one, so I will never write more than 10 words exactly as they appear on any of the 20 websites.

- I will probably also write something that is not correct or add something new because, as I said, my memory is not the best.

Is that fair use? Am I infringing on copyright?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#243
post #242

If I create a news website where I write articles in the following way: - Read 20 different news websites and their story on the same event/topic - Wait an hour, grab a cup of coffee - Sit down to write my article, never from this point I open any of the 20 news websites, I write the story from my head - I don't consult any other source, just write from my memory, and my memory is, let's say, not the best one, so I w…

If you are piece of software then yes.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#244
post #82

The suit demonstrates instances where ChatGTP / Bing Copilot copy from the NYT verbatim. I think it is hard to argue that such copying constitutes "fair use". However, OAI/MS should be able to fix this within the current paradigm: Just learn to recognize and punish plagiarism via RLHF. However, the suit goes far beyond claiming that such copying violates their copyright: "Unauthorized copying of Times Works without p…

> Just learn to recognize and punish plagiarism via RLHF. This is not a RLHF problem. What I was expecting them to do is to keep a bloom filter of ngrams for known copyrighted content, such as enumerating all sets of n=7 consecutive words in an article, and validate against it. The model would only output at maximum n-1 words that look verbatim from the source. But this will blow up in their face. Let's see: - AI com…

I think it is an RLHF problem and that you are right - this will blow up in the faces of the NYT.

Specifically, the NYT examples all seem to be cases where they asked the AI to repeat their articles verbatim? So they ask it to violate copyright and because it's a helpful bot with a good memory, it does so.

Solution: teach the model to refuse requests to repeat articles verbatim. It's easily capable of recognizing when it's being asked to do that. And that's exactly what OpenAI have now done.

So the direct problem the NYT is complaining about - a paywall bypass - is already rectified. Now it would seem to me like the case is quite weak. They could demand OpenAI pay them damages for the time ChatGPT wasn't refusing, but wouldn't they have to prove damages actually happened? It seems unlikely many people used ChatGPT as a paywall bypass for the NYT specifically in the past year. It only knows old articles. OpenAI could be ordered to search their logs for cases where this happened, for example, and then the NYT could be ordered to show their working for the value of displaying a single old article to a non-subscriber, and from that damages could be computed. But it wouldn't be a lot.

That's presumably why the case goes further and argues that OpenAI is in violation even when it isn't repeating text verbatim. That's the only way the NYT can get any significant money out of this situation.

But this case seems much weaker to me. Beyond all the obvious human analogies, there is precedent in the case of search engines where they crawl - and the NYT let them crawl - specifically to enable the creation of a derived data structure. Search engine indexes are understood to be fair use, and they actually do repeat parts of the page verbatim in their snippets. Google once even showed cached versions of whole pages. And browser makers all allow extensions in their stores that strip ads and bypass paywalls, and the NYT hasn't sued them over that either.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#246

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

To me, there is a sense that the news, which is real information about the society that we currently live in, should be availabe to all participants of that society. The notion of being a good citizen requires that one stays informed. Books, movies, videogames etc. don't have that role and are more consumption goods.

It’s a difficult problem with no great answers. If you want news to be free at the point of delivery you want public service news agencies. But that means they’re owned by the government… who are frequently the target of critical reporting.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#247

If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…

There's nothing wrong with scraping openly available data (including data openly available by mistake, as long as you are not aware of it, see the Bluetouff affair).

So the demand to destroy those databases seems very dubious to me.

Of course later violating fair use is another issue.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#248

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

Possibly because once an article is published the author receives no further payment. In all other mediums, there are residuals and royalties to be paid to the creators of the work.

Articles have ads on them, how are they not residual payments based on views?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#249
post #64
post #52

Earlier quoted context omitted.

Sounds like you didn't read the article. Here's a better synoposis: I read a NYT article and publish an exact copy of that article on my website: copyright infringement. Train a model on NYT text and it outputs an exact copy of that text: also copyright infringement.

A small number of outputs of ChatGPT are close enough to training articles to be (probably) copyright infringement. What does that mean? Look up "substantial non-infringing use" and this little court case: https://en.wikipedia.org/wiki/Sony_Corp._of_America_v._Unive... . Now spend a few million on lawyers and roll your dice.

In Sony vs. Universal case, Sony is the producer of a tool where the consumer uses to "time-shift" a broadcast that they legally are allowed to view. Similarly, you can rip your own CDs or photocopy your own books. This case never made reselling those content legal. OpenAI does not train ChatGPT on the content you own - they do it on some undisclosed amount of data that you may or may not have a legal right to access, and then move on and (is shown to) reproduce it nearly verbatim - they may even charge you for the pleasure.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#250
post #60

Earlier quoted context omitted.

Young males that wear Tensorflow branded muscle tank tops and drive Mitsubishi Eclipse convertibles with the vanity plate OVERFIT. They are everywhere these days.

https://i.imgur.com/4tF7q8M.jpg

The text generation is getting quite decent. The limbs disappearing into the car are somewhat less impressive.
Post reply on HN