Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

51–60 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#51

The arguments about being able to mimic New York Times “style” are weak, but the fact that they got it to emit verbatim NY Times content seems bad for OpenAI: > As outlined in the lawsuit, the Times alleges OpenAI and Microsoft’s large language models (LLMs), which power ChatGPT and Copilot, “can generate output that recites Times content verbatim

Sarah Silverman is claiming the same thing about her book.

But I've tried really hard to get ChatGPT to output sentences verbatim from her book and just can't get it to. In fact, I can't even get it to answer simple questions about facts that are in her book but nowhere else -- it just says it doesn't know.

Similarly I haven't been able to reproduce any text in the NYT verbatim unless it's part of a common quote or passage the NYT is itself quoting. Or it's a specific popular quote from an article that went viral, but there aren't that many of those.

Has anyone here ever found a prompt that regurgitates a paragraph of a NYT article, or even a long sentence, that's just regular reporting in a regular article?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#52
post #17

Earlier quoted context omitted.

Take estimated losses of the NYT from this "innovation" and multiply by 10^x where is "x" high enough to make tech companies stop and think before they break laws next time. That would be my approach at least.

which laws are broken exactly? it's not remotely settled law that "training an NN = copyright infringement"

No, we’re seeing the first steps of it (maybe) becoming settled law.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#53
post #16

What are they arguing here? AFAIK reading copyrighted works is not copyright infringement. Copying and selling them is, as the name would suggest, but OpenAI absolutely did not do that. Are they trying to say that LLM training is a special type of reading that should be considered infringement? Seems like a weak case to me. edit: Would be very funny if OpenAI used an educational fair use defense

> AFAIK reading copyrighted works is not copyright infringement. [...] Are they trying to say that LLM training is a special type of reading that should be considered infringement? Nobody can argue that OpenAI was feeding the content to ChatGPT because ChatGPT was bored or was curious about current events. It was fed NYT's content so it would know how to reproduce similar content, for profit. I think getting a case-l…

> It was fed NYT's content so it would know how to reproduce similar content, for profit.

Sounds like journalism school?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#54
post #2

NYT article with a lot more context https://www.nytimes.com/2023/12/27/business/media/new-york-t...

Does it not seem a bit suspect to read the the NYT reporting on their own lawsuit?

The newsroom is a different part of thr company than the legal department. Plus, sometimes your company does something that's newsworthy! Just like all journalism, there's always implicit bias. No reason to get suspicious about a news organization covering the news.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#55
post #12

Can someone explain the technical difference between what search engines do to index newspapers versus what is being claimed here? Is the difference as simple as me being able to get summaries and content from a newspaper from GPT without needing to visit their website?

A search engines principle job is to provide you with links you can find the answer to your question.

The LLMs are ingesting all of that content en masse and would provide you the answer directly, with no compensation to the writers who actually did the research to provide that answer.

Search engines are symbiotic, LLMs are parasitic.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#56
Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States.

And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can do other than put everything you create behind some kind of authentication wall but even then it’s only a matter of time until it leaks anyway.

Pandora’s box is really open, we need to figure out how to live in a world with these systems because it’s an un winnable arms race where only bad actors will benefit from everyone else being neutered by regulation. Especially with the massive pace of open source innovation in this space.

We’re in a “mutually assured destruction” situation now, but instead of bombs the weapon is information.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#57
post #2

NYT article with a lot more context https://www.nytimes.com/2023/12/27/business/media/new-york-t...

Does it not seem a bit suspect to read the the NYT reporting on their own lawsuit?

The NYT also hallucinates from time to time: https://www.nytimes.com/2003/05/11/us/correcting-the-record-...

(That's a story about Jayson Blair, one of their reporters who just plain made up stories and sources for months before getting caught)

Edit: Sheesh, even their apology is paywalled. Wiki background: https://en.wikipedia.org/wiki/Jayson_Blair?wprov=sfla1

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#58
post #21

For me it's quite obvious that if you make a profit from an engine that has as an input copyrighted material, then you owe something to the owner of this copyrighted content. We have seen this same problem with artists claiming stable diffusion engines were using their art.

Do all automakers that now develop electric cars owe Tesla something as they cashed in once they saw Tesla's successful copyrighted material l? A model is semantic, it contains the idea which is not copyrightable. Only how it is expressed could be copyrighted (i.e. if it outputs the copyright work verbatim). If this were not the case we would have plenty of monopolies and the world would fall apart.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#59

Does anyone know what the copyright status of LLM generated content is? That is, if I feed a NYT article into GPT4 and say, summarize this article, and then publish that summary, is there argument or precedent that says that is or is not copyright infringement? Asking for a friend.

Technically, you just send a request to OpenAI and they are the ones who feed it into GPT4. Although I'd argue this is irrelevant to your question, the law works in mysterious way so perhaps it carries some importance.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#60
This is just rent seeking from dying media instead of working on creating something new in my view.

AI indeed is reading and using material sa a source, but is deriving results based on that material. I think this should be allowed, but now it is a fight who has better paid politicians pretty much.

I am open to hear other thoughts.

Post reply on HN