Live data from Hacker News

New York Times considers legal action against OpenAI as copyright tensions swirl

npr.org

41–50 of 383 posts

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#41

Earlier quoted context omitted.

Did the Times grant a license to every router on the internet to transmit its intellectual property to other routers? If not, the judge should grant an injunction contingent on requiring the Times to verify that every person who accesses their content is doing so only over routers and other devices with express written authorization, for every step in the process. Maybe even extend it to browsers and client libraries…

OpenAI is pretty clearly using their work to make derivative content that in certain cases (CNET) is a direct competitor. Honestly, this seems open and shut

> to make derivative content

How would they prove this? Is it safe to say that each article used has a nearly meaningless influence on the weights?

Could this be used as a defense? Perhaps train a (smaller) model, remove a single article, and show how it doesn't influence performance?

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#42

Earlier quoted context omitted.

Did the Times grant a license to every router on the internet to transmit its intellectual property to other routers? If not, the judge should grant an injunction contingent on requiring the Times to verify that every person who accesses their content is doing so only over routers and other devices with express written authorization, for every step in the process. Maybe even extend it to browsers and client libraries…

OpenAI is pretty clearly using their work to make derivative content that in certain cases (CNET) is a direct competitor. Honestly, this seems open and shut

It’s not making a derivative product though. The fact that they compete or not doesn’t matter for copyright.

It’s not like it’s ok to violate copyright if you don’t compete. It’s still illegal to take a NYT article and print it on a t-shirt.

The issue is that copyright law doesn’t prevent the kind of model training as there’s no clearly derived work. I don’t think that’s been tested in courts yet, but I expect it won’t be found to be copyright because there’s other precedent that influenced is not infringement.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#43

Earlier quoted context omitted.

Did the Times grant a license to every router on the internet to transmit its intellectual property to other routers? If not, the judge should grant an injunction contingent on requiring the Times to verify that every person who accesses their content is doing so only over routers and other devices with express written authorization, for every step in the process. Maybe even extend it to browsers and client libraries…

OpenAI is pretty clearly using their work to make derivative content that in certain cases (CNET) is a direct competitor. Honestly, this seems open and shut

No more open and shut than the NYT having the right to sue people writing editorial news stories with no new content based on reading their news (along with many other sources).

This issue is ultimately going to come down to the transformative clause of fair use. The fact is that the _model_ is unquestionably a transformative product of the inputs, and a judge ruling otherwise is going to cause a cascading shitstorm of litigation and put a chill through the creative economy. The outputs of the model under certain conditions can be guided towards copyright infringement, and any sane ruling will focus on protecting rightsholders from overly derivative model outputs. In all likelihood the precedent will be that the standard for being transformative will be raised for "algorithmically generated" content, and the people who distribute that content will still be fully liable in the event of infringement, with "I didn't know, the AI did it" not being an acceptable defense.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#44
post #26

Earlier quoted context omitted.

Right. But the comment above said the issue is with the copy made for training. The "read it and memorized it" copy, not the "write it back down" copy.

That copy didn’t go into a human’s brain. It went into GPU memory. It’s a copy under the law, no different from copying a Taylor Swift mp3 onto a flash drive. Whether that copy was fair use is the key question.

If I buy a license to listen to a Taylor Swift mp3, I can copy it onto a flash drive (or anywhere I like to use it). And that’s fine until I distribute copies to others.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#45
post #4
post #2

Interesting point… > A top concern for The Times is that ChatGPT is, in a sense, becoming a direct competitor with the paper by creating text that answers questions based on the original reporting and writing of the paper's staff.

It’s sort of ironic though since the Times is “news” where as GPT is built on historic texts. Of course that gap will tighten until we have near real-time models. But that’s not the reality today.

I think it’s ironic because news organizations observe the world and then synthesize it into reporting.

Imagine if someone doing a thing sued NYT for watching them do it, linking it to other issues and producing a new article.

News itself is a derived content that’s dependent on other people doing things.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#46
>If, when someone searches online, they are served a paragraph-long answer from an AI tool that refashions reporting from The Times, the need to visit the publisher's website is greatly diminished, said one person involved in the talks.

If, when someone reads a newspaper, they are served a paragraph-long answer from an NYTimes reporter that refashions reporting from local sources, the need to interact with the local sources is greatly diminished.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#47
post #5
post #4

Earlier quoted context omitted.

It’s sort of ironic though since the Times is “news” where as GPT is built on historic texts. Of course that gap will tighten until we have near real-time models. But that’s not the reality today.

The Times is not merely news, it is the paper of record for the united states (meaning its historical articles are indexed and commonly used to establish prior facts).

[deleted]

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#48

Earlier quoted context omitted.

Fair use only covers very limited circumstances which probably does not include selling a subscription (ChatGPT+). If you’re selling a repackaged reproduction of someone else’s copyrighted works, that’s never protected by fair use.

But you generally can't copyright the underlying facts, only the creative expression in the article/work itself. If they can get it to extract just the facts and how they relate to one another, that wouldn't really be protected by copyright. They might try to bring back the "Hot News" doctrine though.

The models not only were trained on copyrighted works, but when prompted will substantially and repeatedly reproduce copies of creative work (not just facts). That's infringement. https://www.theverge.com/2023/7/9/23788741/sarah-silverman-o...

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#49
I don't think it's an exaggeration to say that LLMs might lead to the end of the open web, or at least a drastically reduced version of it. So much of these model's utility is in directly competing with the producers of the training data. Content creators and aggregators are seeing more and more reason to restrict and limit access, to avoid having AI companies consume all of their data and then be the ones making money from it going forward.

I fear that LLMs are going to cause the internet to be a much worse and less open space.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#50

Earlier quoted context omitted.

Fair use only covers very limited circumstances which probably does not include selling a subscription (ChatGPT+). If you’re selling a repackaged reproduction of someone else’s copyrighted works, that’s never protected by fair use.

But you generally can't copyright the underlying facts, only the creative expression in the article/work itself. If they can get it to extract just the facts and how they relate to one another, that wouldn't really be protected by copyright. They might try to bring back the "Hot News" doctrine though.

My naive reading of it sees that it only applies for time-sensitive facts. Considering the training times required, and knowledge cutoffs, could this apply?
Post reply on HN