Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

61–70 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#61

Earlier quoted context omitted.

The article mentions that ChatGPT will absolutely parrot back NYTimes article text verbatim. So yes, it's copyright infringement.

Sections of this statement absolutely parrot back NYTimes article text vebatim depending how you look at it. What's the line? 3 sequential verbatim words? 5? 8?

We'll find out won't we ;)

You have to imagine these limits are already fairly known within the legal community... If you're accused of copying/republishing my published work there will be some minimal threshold of similarity I would need to prove in order to seek damages.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#62
post #50

Time to move LLM training to Japan who passed a law giving free reign to train LLMs on copyrighted material.

And along with it, a different ideology would tag along.

Not a bad thing, but Japan or China or Russia, don’t align with Anglo centered ideology, so keep that in mind.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#63

Earlier quoted context omitted.

(2023 - 1851) * 365 * 200 = 12,556,000 Yep, so a few million ripped off articles is plausible.

Everything from 1851 to 1927 ought to be in the public domain, though. If the goal of training an AI is just "to mimic a style" there are absolutely humongous amounts of text that's totally free of any copyright restrictions.

Yes, there is large amounts of public domain text available, but does anyone believe this is a restriction that was imposed when feeding the models?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#64

The arguments about being able to mimic New York Times “style” are weak, but the fact that they got it to emit verbatim NY Times content seems bad for OpenAI: > As outlined in the lawsuit, the Times alleges OpenAI and Microsoft’s large language models (LLMs), which power ChatGPT and Copilot, “can generate output that recites Times content verbatim

Sarah Silverman is claiming the same thing about her book. But I've tried really hard to get ChatGPT to output sentences verbatim from her book and just can't get it to. In fact, I can't even get it to answer simple questions about facts that are in her book but nowhere else -- it just says it doesn't know. Similarly I haven't been able to reproduce any text in the NYT verbatim unless it's part of a common quote or p…

The complaint has specific examples they got from ChatGPT.

There is a precedent: There were some exploit prompts that could be used to get ChatGPT to emit random training set data. It would emit repeated words or gibberish that then spontaneously converged on to snippets of training data.

OpenAI quickly worked to patch those and, presumably, invested energy into preventing it from emitting verbatim training data.

It wasn’t as simple as asking it to emit verbatim articles, IIRC. It was more about it accidentally emitting segments of training data for specific sequences that were semi rare enough.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#65

The arguments about being able to mimic New York Times “style” are weak, but the fact that they got it to emit verbatim NY Times content seems bad for OpenAI: > As outlined in the lawsuit, the Times alleges OpenAI and Microsoft’s large language models (LLMs), which power ChatGPT and Copilot, “can generate output that recites Times content verbatim

I can get a printer to emit verbatim NYT content, and with a lot less effort than getting it out of an LLM. I find this capability of infringement equals infringement argument incredibly weak.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#66

Does anyone know what the copyright status of LLM generated content is? That is, if I feed a NYT article into GPT4 and say, summarize this article, and then publish that summary, is there argument or precedent that says that is or is not copyright infringement? Asking for a friend.

No one knows, this is new territory. Maybe the fermi filter is litigating an AI that would otherwise save humanity.

Or the filter could be the other way, failing to litigate an AI to decelerate it's progress, and a risk of augmenting the underlying society too much too quickly.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#67

What are they arguing here? AFAIK reading copyrighted works is not copyright infringement. Copying and selling them is, as the name would suggest, but OpenAI absolutely did not do that. Are they trying to say that LLM training is a special type of reading that should be considered infringement? Seems like a weak case to me. edit: Would be very funny if OpenAI used an educational fair use defense

It should be noted that there are explicit exemptions to allow copying program data intro RAM and into CPU registers (in many licenses). Whether that is truly necessary or not is at best debatable, but arguably training a model (especially one you then distribute or give access to) on copyrighted data is vastly different from regular copying into memory and should require explicit licensing.

The fact that the model can reproduce large chunks of the original text verbatim is proof positive that it contains copies of the original text encoded in its weights. If I wrote a program that crawled the NYT site, zipping the contents, and retrieved articles based on keyword searches and made them available online, would you not say I'm infringing their copyright?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#68

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

I don't think they're looking to prevent the inevitable, but rather see a target with a fat wallet from which a lot of money can be extracted. I'm not saying this in a negative way, but much of the "this is outrageous!" reaction to AI hasn't been about the building of models, but rather the realization that a few players are arguably getting very rich on those models so other people want their piece of the action.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#69

Earlier quoted context omitted.

Does it not seem a bit suspect to read the the NYT reporting on their own lawsuit?

The newsroom is a different part of thr company than the legal department. Plus, sometimes your company does something that's newsworthy! Just like all journalism, there's always implicit bias. No reason to get suspicious about a news organization covering the news.

If Apple is in a lawsuit I'm not going to go to the Apple media relations page for the story. What about the NYT, also a for-profit company, makes it more principled than Apple, other than that they say they are?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#70
post #2

NYT article with a lot more context https://www.nytimes.com/2023/12/27/business/media/new-york-t...

[edit: they have since opened a comment section to the article.] It is unfortunate that the NYTimes don’t allow reader comments to this article. I like some of the NYTimes content, but in this case use of chatGPT is infinitely more valuable to me than subscribing to the NYTimes, so I would like to explain this concept and the associated risks by their litigation without cancelling my subscription.

Maybe it is time to move training of models to Japan that has explicitly adapted AI friendly legislation that allows training on previously copyrighted materials. My best guess is that if the inputs were legally obtained, then the output doesn’t violate anything until someone publishes it. Similar to how reading a newspaper in a public library is legal but copying its content verbatim and republishing is not.

Post reply on HN