Live data from Hacker News

NY Times copyright suit wants OpenAI to delete all GPT instances

arstechnica.com

481–490 of 921 posts

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#481
post #419

The NYT is preparing for a tsunami by building a sandcastle. Big picture, this suit won’t matter, for so many reasons. To enumerate a few: 1. Next gen LLMs will be trained exclusively on “synthetic”/public data. GPT-4V can easily whitewash its entire copyrighted training corpus to be unrecognizably distinct (say reworded by 40%, authors/sources stripped, etc). Ergo there will be no copyright material for GPT-5 to reg…

I'm sorry but this is such a bad take. Nice appeal to consequences. In my view, the New York Times is entirely justified in pursuing legal action. They invested time and effort in creating content, only to have it used without permission for monetary gain. A clear violation. Analyzing the factors involved for a "fair use" consideration: Purpose and Character of the Use: While the argument for transformation might hol…

> it's clearly not helping their market value if people are checking on ChatGPT instead of reading a NYT article.

People are not using ChatGPT as a replacement for current news, and because of hallucinations, no one should be using it for past news either. I wouldn't remotely call ChatGPT a competitor of NYT traffic, like I would Reuters or other news outlets.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#482

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

A book, TV show, movie, video game, album, or comic book is not available on the internet served by the copyright holder’s own servers with no authentication or authorization checks. But the NYT is available in that way.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#483

Earlier quoted context omitted.

Ok but it's not

Definition of Transformative Use: The legal concept of transformative use involves significantly altering the original work to create new expressions, meanings, or messages. AI models like GPT don't merely reproduce text; they analyze, interpret, and recombine information to generate unique responses. This process can be argued as creating new meaning or purpose, different from the original works. In the case of the…

Nope, it doesn't work that way. The fact that the LLM can regurgitate original articles doesn't remove the possibility that training can be considered transformative work, or more in general that using copyrighted material for training can be considered fair use.

Rather, verbatim reproduction is the proof that copyrighted materials was used. Then the court has to evaluate whether it was fair use. Without verbatim reproduction, the court might just say that there is not enough proof that the Times's work was important for the training, and dismiss the lawsuit right away.

Instead, the jury or court now will almost certainly have to evaluate OpenAI's operation against the four factors.

In fact, I agree with the parent that ingesting text and creating a representation that can critique historical facts using material that came from the Times is transformative. An LLM is not just a set of compressed texts, people have shown for example that some neurons fire when you are talking of specific historical periods or locations on Earth.

However, I don't think that the trasformative character is enough to override the other factors, and therefore in the end it won't/shouldn't be considered fair use IMHO.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#484
post #482

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

A book, TV show, movie, video game, album, or comic book is not available on the internet served by the copyright holder’s own servers with no authentication or authorization checks. But the NYT is available in that way .

But some are? I believe The Atlantic and The Economist are hard paywalled.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#485

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

We are also happy to use open source, yet what open source alternatives are there for news that don't get shot down by the media or besmirched?

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#486
post #436

Earlier quoted context omitted.

Largely because "news" aka facts is not and should not be copyrightable, so while the style, and exact format of the article may be copyrightable, the facts contained within are not. This makes a news story copyright murky in the eyes of wider society unlike a clearly 100% creative work like a TV Show or Movie. Further the news themselves self cannibalize, how many stories are just rewrites of stories from other outl…

why it is OK for the Washington Post to copy the NY times, but not ok for OpenAI or Archive.org? If the Washington Post printed an article from the NY Times nearly verbatim and without attribution, it would not be OK and surely they would take legal action.

Yes, because The NY Times is copyrighting the body of work. They are not copyrighting the "facts" themselves but the distillation of these facts into a body of work. Anyone is free to take the facts and produce their own works but not to lift the body of work verbatim that the NY Times created (plagiarize).

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#487

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

I wouldn't say OpenAI has exactly the same attitude, since they also pulled in thousands of books. Their position has been that it's not piracy, since they don't republish the books; effectively the AI just reads them and learns from them. If GPT can be made to reproduce the original articles, that's a more difficult argument to make.

It turns out you can reproduce articles with next-token prediction when the articles are quoted all over the dataset.

The articles themselves are indisputably not a part of the model, because it doesn't store text at all. OpenAI's position is correct; people just underestimated how well the AI learns from reading, especially when it reads the same text in a bunch of different places because it's being quoted/excerpted.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#488
post #418

People who think the examples the lawsuit are “fair use” need to consider what that would mean. We’re basically going to let a few companies consolidate all the value on the Internet into their black boxes with basically no rules … that seems very dangerous to me. I hope a court establishes some rules of engagement here, even if it’s not this case.

a court has established this already in japan, where they said anything goes for ai so its best to not to lose a competitive edge with things that people openly publish on the internet, if you put it out there for everyone to see then expect other people to use it

A court in Japan will have no impact on the outcome of a copyright lawsuit in USA. Not to mention that it doesn't really matter how a Japanese court ruled since it's all governed by treaties anyway. They will change their laws if required to.

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#489

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

Historically newspapers leaned more on competition law than copyright, because their pages are supposed to be filled with non-copyrightable facts.[1] Copying part, but not all, of a factual article, significantly after the relevant event, was considered to be a promotion (not unfair competition) and a nice thing to do for the journalists. Things change, people lose sight of the original principles. [1] https://en.m.w…

These days most news is mixed with analysis [1] (which is often biased). I wonder if part of the reason for this shift is that analysis is copyrightable. It also seems like the number of opinion articles is ever expanding [2], though I don't have any hard numbers on that.

[1]: https://guides.library.cornell.edu/evaluate_news/source_bias

[2]: https://www.newsmediaalliance.org/rise-of-opinion-section/ Interestingly there's a banner at the top of that link touting an agreement between Axel Springer and OpenAI.

EDIT: formatting

Re: NY Times copyright suit wants OpenAI to delete all GPT instances

#490

It's interesting to me the ambiguous attitude people have to reproducing news content. Whenever there is a story from NYT on HN (or any other large media outlet), the top comment is almost always a link to an archived version which reproduces the text verbatim. And this seems to be tolerated as the norm. And yet, whenever there is a submission about a book, a TV show, a movie, a video game, an album, a comic book, or…

I wouldn't say OpenAI has exactly the same attitude, since they also pulled in thousands of books. Their position has been that it's not piracy, since they don't republish the books; effectively the AI just reads them and learns from them. If GPT can be made to reproduce the original articles, that's a more difficult argument to make.

I can understand an argument about the AI needing to know basic history. News is just how we report history in the making, but it's not generally accepted as solid until some time after the events when we can get more context.

Isn't this what the Associated Press is intended for, a stream of news trying to report just the facts and happenings of the day? That's quite a bit different than a NYT article intending to inform but also convince someone of a position of some sort.

Feeding an AI opinionated news compared to "just the facts, ma'am" seems risky from a bias perspective.

Post reply on HN