It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…
NY Times copyright suit wants OpenAI to delete all GPT instances
241–250 of 921 posts
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#242- Read 20 different news websites and their story on the same event/topic
- Wait an hour, grab a cup of coffee
- Sit down to write my article, never from this point I open any of the 20 news websites, I write the story from my head
- I don't consult any other source, just write from my memory, and my memory is, let's say, not the best one, so I will never write more than 10 words exactly as they appear on any of the 20 websites.
- I will probably also write something that is not correct or add something new because, as I said, my memory is not the best.
Is that fair use? Am I infringing on copyright?
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#243If I create a news website where I write articles in the following way: - Read 20 different news websites and their story on the same event/topic - Wait an hour, grab a cup of coffee - Sit down to write my article, never from this point I open any of the 20 news websites, I write the story from my head - I don't consult any other source, just write from my memory, and my memory is, let's say, not the best one, so I w…
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#244The suit demonstrates instances where ChatGTP / Bing Copilot copy from the NYT verbatim. I think it is hard to argue that such copying constitutes "fair use". However, OAI/MS should be able to fix this within the current paradigm: Just learn to recognize and punish plagiarism via RLHF. However, the suit goes far beyond claiming that such copying violates their copyright: "Unauthorized copying of Times Works without p…
> Just learn to recognize and punish plagiarism via RLHF. This is not a RLHF problem. What I was expecting them to do is to keep a bloom filter of ngrams for known copyrighted content, such as enumerating all sets of n=7 consecutive words in an article, and validate against it. The model would only output at maximum n-1 words that look verbatim from the source. But this will blow up in their face. Let's see: - AI com…
Specifically, the NYT examples all seem to be cases where they asked the AI to repeat their articles verbatim? So they ask it to violate copyright and because it's a helpful bot with a good memory, it does so.
Solution: teach the model to refuse requests to repeat articles verbatim. It's easily capable of recognizing when it's being asked to do that. And that's exactly what OpenAI have now done.
So the direct problem the NYT is complaining about - a paywall bypass - is already rectified. Now it would seem to me like the case is quite weak. They could demand OpenAI pay them damages for the time ChatGPT wasn't refusing, but wouldn't they have to prove damages actually happened? It seems unlikely many people used ChatGPT as a paywall bypass for the NYT specifically in the past year. It only knows old articles. OpenAI could be ordered to search their logs for cases where this happened, for example, and then the NYT could be ordered to show their working for the value of displaying a single old article to a non-subscriber, and from that damages could be computed. But it wouldn't be a lot.
That's presumably why the case goes further and argues that OpenAI is in violation even when it isn't repeating text verbatim. That's the only way the NYT can get any significant money out of this situation.
But this case seems much weaker to me. Beyond all the obvious human analogies, there is precedent in the case of search engines where they crawl - and the NYT let them crawl - specifically to enable the creation of a derived data structure. Search engine indexes are understood to be fair use, and they actually do repeat parts of the page verbatim in their snippets. Google once even showed cached versions of whole pages. And browser makers all allow extensions in their stores that strip ads and bypass paywalls, and the NYT hasn't sued them over that either.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#245Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#246It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…
To me, there is a sense that the news, which is real information about the society that we currently live in, should be availabe to all participants of that society. The notion of being a good citizen requires that one stays informed. Books, movies, videogames etc. don't have that role and are more consumption goods.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#247If you forget about the LLM aspect, and simply build a product out of (legally) scraped NYT articles, is that fair use? Let's say I host these, offer some indexing on it, and rewrite articles. Something like, summarise all articles on US-UK relationships over past 5 years. I charge money for it, and all I pay NYT is a monthly subscription fee. To keep things simple, let's say I never regurgitate chunks of verbatim NY…
So the demand to destroy those databases seems very dubious to me.
Of course later violating fair use is another issue.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#248It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…
Possibly because once an article is published the author receives no further payment. In all other mediums, there are residuals and royalties to be paid to the creators of the work.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#249Earlier quoted context omitted.
Sounds like you didn't read the article. Here's a better synoposis: I read a NYT article and publish an exact copy of that article on my website: copyright infringement. Train a model on NYT text and it outputs an exact copy of that text: also copyright infringement.
A small number of outputs of ChatGPT are close enough to training articles to be (probably) copyright infringement. What does that mean? Look up "substantial non-infringing use" and this little court case: https://en.wikipedia.org/wiki/Sony_Corp._of_America_v._Unive... . Now spend a few million on lawyers and roll your dice.
Re: NY Times copyright suit wants OpenAI to delete all GPT instances
#250Earlier quoted context omitted.
Young males that wear Tensorflow branded muscle tank tops and drive Mitsubishi Eclipse convertibles with the vanity plate OVERFIT. They are everywhere these days.
https://i.imgur.com/4tF7q8M.jpg