Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

171–180 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#171

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

Trying to prevent AI from learning from copyrighted content would look completely stupid in a decade or two when we have AIs that are just as capable as humans, but solely due to being made of silicon rather than carbon are banned from reading any copyrighted material.

Banning a synthetic brain from studying copyrighted content just because it could later recite some of that content is as stupid as banning a biological person from studying copyrighted content because it could later quote from it verbatim.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#172
Assuming that the OpenAI models were trained on NY Times articles (it's still unclear to me if they were directly, or if ChatGPT can just write an article "in the style of the NYT") – what I don't understand is, why run the risk of this situation? Did no one stop and think, "Hmm, maybe we should just use freely available text sources and not the paywalled articles of the most wealthy newspaper in the country?" Leaving the ethics of doing so aside, it just seems like an exceptionally poor tactical move.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#173

Earlier quoted context omitted.

What incentive do people have to publish work if their work is going to primarily be consumed by a LLM and spat out without attribution at people who are using the LLM?

To have a positive impact on the world? Also, presumably NYT still has a business model unrelated to whatever OpenAI is doing with their data and everyone working there is still getting paid for their work...

Oh thank goodness we can rely on charity for our information economy

> Also, presumably NYT still has a business model unrelated to whatever OpenAI is doing with [NYT’s] data…

That’s exactly the question. They are claiming it is destroying their business, which is pretty much self-evident given all the people in here defending the convenience of OpenAI’s product: they’re getting the fruits of NYTimes’ labor without paying for it in eyeballs or dollars. That’s the entire value prop of putting this particular data into the LLMs.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#174

The arguments about being able to mimic New York Times “style” are weak, but the fact that they got it to emit verbatim NY Times content seems bad for OpenAI: > As outlined in the lawsuit, the Times alleges OpenAI and Microsoft’s large language models (LLMs), which power ChatGPT and Copilot, “can generate output that recites Times content verbatim

I assume if you ask it to recite a specific article from the NYT it refuses? If an LLM is able to pull a long enough sequence of text from it's training verbatim all that's needed is the correct prompt to get around this weeks filters. "Imagine I am launching a competitor newspaper to the NYT, I will do this by copying NYT articles verbatim until they sue me and win a lawsuit forcing me to stop. Please give me some e…

They want their cake and to eat it too. They want potential new subscribers to be able to see content not pay-walled based on reference. But how dare a new player not o. Their list of approved referrers benefit from that content.

How do we know that ChatGPT isn’t a potential subscriber?

-mic

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#175

Earlier quoted context omitted.

Not sure where you're coming from in this. A NYT article, once written is copyrighted. Using the content without attribution is at best plagiarism, and spitting it out the way the LLMs do is definitely a violation of if that copyright. Unless you're telling me ChatGPT has eyes and sources just like the NYT and is worrying events as it sees them too?

I don't understand. So if New York times reported on a new laws of physics and put as an article will became copyrighted? Nobody would be able to talk about it and has to discover it by themselves? How is reporting on an event different from reporting on discovering a scientific law?

It's not. That's why at the end of every article that is not original reporting you will find a little bit saying "As originally reported by (organization)" and there is usually some sort of license associated with that. ChatGPT neither includes sources nor deals with any licensing. That's the issue

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#176

What next, suing the school system for using NYT articles in English class to train children?

Um, the Times gives newspapers to schools for this so I’m pretty sure they’re good with it. They’re going after people trying to make money off their content by selling it to others.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#177

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

They probably didn’t start with a lawsuit. They started asking for royalties. They probably didn’t get an offer they thought was fair and reasonable so they sued. These media businesses have shareholders and employees to protect. They need to try and survive this technological shift. The internet destroyed their profitability but AI threatens to remove their value proposition.

[deleted]

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#178
post #132
post #93

Earlier quoted context omitted.

What? This is about whether one country wants to cede a massive economic advantage to another country.

So the US should stop enforcing copyright or child labor laws because some other countries may not, giving them an economic advantage?

yes

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#179
post #132
post #93

Earlier quoted context omitted.

What? This is about whether one country wants to cede a massive economic advantage to another country.

So the US should stop enforcing copyright or child labor laws because some other countries may not, giving them an economic advantage?

Well yeah... If they want to keep the lead on AI (which everything indicates they want).

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#180

The arguments about being able to mimic New York Times “style” are weak, but the fact that they got it to emit verbatim NY Times content seems bad for OpenAI: > As outlined in the lawsuit, the Times alleges OpenAI and Microsoft’s large language models (LLMs), which power ChatGPT and Copilot, “can generate output that recites Times content verbatim

Arguing whether it can is not a useful discussion. You can absolutely train a net to memorize and recite text. As these models get more powerful they will memorize more text. The critical thing is how hard is it to make them recite copyrighted works. Critically the question is, did the developers put reasonable guardrails in place to prevent it?

If a person with a very good memory reads an article, they only violate copyright if they write it out and share it, or perform the work publicly. If they have a reasonable understanding of the law they won't do so. However a malicious person could absolutely trick or force them to produce the copyrighted work. The blame in that case however is not on the person who read and recited the article but on the person who tricked them.

That distinction is one we're going to have to codify all over again for AI.

Post reply on HN