Live data from Hacker News

New York Times considers legal action against OpenAI as copyright tensions swirl

npr.org

121–130 of 383 posts

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#121
post #54
post #36

Earlier quoted context omitted.

Do you think OpenAI doesn’t subscribe? I too load creative works into all sorts of temporary structures in order to read the paper. I don’t need to license it, I pay for a subscription. Should I pay more if I memorize the paper? Should I pay more if I read it to my sick friend in the hospital? Should I pay more if I save copies to my own hard drive and grep for words in the files? Should robots have a higher subscrip…

Are you using this knowledge to produce a product that materially undercuts NYT revenue?

That doesn’t matter though. If I licensed material why should it count if I compete or not.

If I watch Lebron James play and use it to develop an athletic training program, does it matter if I play baseball? Or WNBA? Or does it only matter if I play against him in the championship?

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#124

>If, when someone searches online, they are served a paragraph-long answer from an AI tool that refashions reporting from The Times, the need to visit the publisher's website is greatly diminished, said one person involved in the talks. If, when someone reads a newspaper, they are served a paragraph-long answer from an NYTimes reporter that refashions reporting from local sources, the need to interact with the local…

So? The local source is free to sue the NYT for copyright infringement if they so wish.

This certainly happens in the UK with different newspapers, and starting with a letter demanding payment, but the same idea.

Very few people bother doing this for much the same reason very few bother fighting any of the other terrible decisions made by corporations with legal departments whose annual cost exceeds their personal lifetime earnings.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#125

Earlier quoted context omitted.

> we end up with an inferior product cannibalizing a superior one and driving it out of business. In case that print is meant by inferior product: The same argument could've been brought up for Napster, where traditional distribution via CD printing through music labels are the inferior product driving the superior one out of business. Or rather it's big labels suing Napster out of business. I also hold a dislike for…

For me the question is highly debatable. As far as I'm aware the training of AI works by crawling various content from the Internet, so they're using a product, which are NY articles to train an AI, meaning they're using their content to help creating a product. But in this sense, shouldn't they be suing Google as well? Since Google as a search engine, also crawls the web and shows their articles in their search resu…

News organizations in other jurisdictions already have achieved settlements with Google (which has much deeper pockets than OpenAI)

But there's a fairly obvious difference in use between using content to index it and point to it and generate revenue for it and using content to generate alternative content...

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#126

Earlier quoted context omitted.

> It’s not even a question of “fair use” because OpenAI isn’t providing anybody a copy “Fair Use” can apply to essentially any of the exclusive rights under copyright, including making (with or without distributing) a copy or derivative work. > Probability models really just aren’t copyrightable to begin with. “Probability models” that are built from copyrightable works either: (1) have a sufficient human creative in…

But OpenAI is neither copying nor deriving a work. Style (which can be described as a probability model) is not copyrightable. You're halfway there with #2. The output is not copyrightable, but unless you can actually point to a sequence of words from the original it can't be infringing.

> But OpenAI is neither copying nor deriving a work.

They absolutely are copying in the course of training the model, and they are doing something that often looks a lot like copying when producing output with the model.

> Style (which can be described as a probability model) is not copyrightable.

Style is not all that can be described in an LLM's "probability model", otherwise LLM models would never be able to reproduce content from the training set, which they do.

> The output is not copyrightable, but unless you can actually point to a sequence of words from the original it can’t be infringing.

While verbatim copying (“a sequence of words from the original”) is infringing, mechanical copying with a mechanical substitution filter (even a wickedly complex one) is also copying, and violates copyright.

And ChatGPT is quite capable of producing verbatim text from its training set.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#128

Don't humans operate similarly? We gain knowledge through experiences. These AI models effectively condense a vast amount of experience data into weights. Considering the global race in AI advancements, I'm skeptical about the success of these copyright claims. I do find it hypocritical that OpenAI says that other LLMs can't be trained on data generated by their LLMs.

True

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#130

Earlier quoted context omitted.

For me the question is highly debatable. As far as I'm aware the training of AI works by crawling various content from the Internet, so they're using a product, which are NY articles to train an AI, meaning they're using their content to help creating a product. But in this sense, shouldn't they be suing Google as well? Since Google as a search engine, also crawls the web and shows their articles in their search resu…

News organizations in other jurisdictions already have achieved settlements with Google (which has much deeper pockets than OpenAI) But there's a fairly obvious difference in use between using content to index it and point to it and generate revenue for it and using content to generate alternative content...

If you're using their content to generate more content, doesn't it fall under fair use?
Post reply on HN