Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

291–300 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#291

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

SciHub was an early warning, IMHO, that there's a strong risk of the first world fumbling the ball so badly with IP that tech ecosystems start growing in the third world instead. The dominant platform for distributing scientific journal papers is no longer Western. Maybe SciHub is economically inconsequential, but LLM's certainly are not!

Imagine if California had banned Google spidering websites without consent, in the late 90's. On some backwards-looking, moralizing "intellectual property" theory, like the current one targeting LLM's. 2/3rd of modern Silicon Valley wouldn't exist today, and equivalent ecosystems would have instead grown up in, who knows where. Not-California.

We're all stupidly rich and we have forgotten why we're rich in the first place.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#292

Here come the innovation sponges. If this goes through then the models that the general public have access are going to be severely neutered while the ownership class will have a much better model that will never see the light of day due to legal risks and claims like this - therefore increasing the disparity between us all.

[flagged]

I’d like a productive comment if you have it.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#293

The arguments about being able to mimic New York Times “style” are weak, but the fact that they got it to emit verbatim NY Times content seems bad for OpenAI: > As outlined in the lawsuit, the Times alleges OpenAI and Microsoft’s large language models (LLMs), which power ChatGPT and Copilot, “can generate output that recites Times content verbatim

Arguing whether it can is not a useful discussion. You can absolutely train a net to memorize and recite text. As these models get more powerful they will memorize more text. The critical thing is how hard is it to make them recite copyrighted works. Critically the question is, did the developers put reasonable guardrails in place to prevent it? If a person with a very good memory reads an article, they only violate…

> Critically the question is, did the developers put reasonable guardrails in place to prevent it?

Why? If I steal a bunch of unique works of art and store them in my house for only me to see, am I still committing a crime?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#294
post #99
post #21

For me it's quite obvious that if you make a profit from an engine that has as an input copyrighted material, then you owe something to the owner of this copyrighted content. We have seen this same problem with artists claiming stable diffusion engines were using their art.

If you study copyrighted material for four years at a university and then go on to earn money based on your education, do you owe something to the authors of your text books? I'm not sure how we should treat LLMs with respect to publicly accessible but copyrighted material, but it seems clear to me that "profiting" from copyrighted material isn't a sufficient criteria to cause me to "owe something to the owner".

Do people ever get tired of this argument that relies on anthropomorphizing these AI black boxes?

A computer isn't a human, and we already have laws that have a different effect depending on if it's a computer doing it or a human. LLMs are no different, no matter how catchy hyping them up as being == Humans may be.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#295
post #101

The way to view this kind of parasitism is how we look at patent trolls. When you look at the RIAA/MPAA lawsuits, while I don't agree with them, at least file sharing was basically a canonical form of copyright infringement. With LLMs we have an aspect of a text corpus that the creators were not using (the language patterns) and had no plans for or even idea that it could be used, and then when someone comes along an…

But it’s theirs, they created it and should therefore benefit from it. I’m honestly shocked at how much these companies are getting away with. It’s piracy on a massive scale. You can get a little discombobulated reading the comments from the nerds / subject idiots on this site.

George R. R. Martin authored A Game of Thrones, but lost in-court against Google when Google Books reproduced parts of his text verbatim: https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....

No piracy or even AI was required, here. Google's defense was that their product couldn't reproduce the book in it's entirety, which was proven and made the prosecution about Fair Use instead. Given that it was much harder to prosecute on those grounds, Google tried coercing the authors into a settlement before eventually the District Court dropped the case in Google's favor altogether.

OpenAI's lawyers are aware of the precedent on copyright law. They're going to argue their application is Fair Use, and they might get away with it.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#296
post #274
post #133

Earlier quoted context omitted.

The word ‘moot’ does not mean what you think it means.

It can do though. While the proper definition is "worthy of discussion / debatable", it can also refer to a pointless debate. "Moot derives from gemōt, an Old English name for a judicial court. Originally, moot referred to either the court itself or an argument that might be debated by one. By the 16th century, the legal role of judicial moots had diminished, and the only remnant of them were moot courts, academic mo…

Do you really think the commenter meant to use moot to mean “purely academic?”

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#297

Surprised they don't mention Bard anywhere in the article. I wonder if the NYT has worked out some sort of licensing deal with Google for Bard, or if Bard isn't trained on NYT data? The lawsuit mentions this, so maybe they did work out some agreement to license their data: "For months, The Times has attempted to reach a negotiated agreement with Defendants, in accordance with its history of working productively with…

Maybe Bard's apparent "behindness" is less about Google's technical merits or lack thereof, and more about it being built with a sense of legal maturity that the competitors don't yet have. After all, Google must have some experience in this space, and we've seen them simply refuse to deploy Bard in regions where (presumably) there is too much legal uncertainty. If 2024's Gemini performs similarly to GPT4 while also navigating legal landmines, maybe it comes out ahead.

Or maybe Bard's lawsuit just hasn't come yet.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#298
post #293

Earlier quoted context omitted.

Arguing whether it can is not a useful discussion. You can absolutely train a net to memorize and recite text. As these models get more powerful they will memorize more text. The critical thing is how hard is it to make them recite copyrighted works. Critically the question is, did the developers put reasonable guardrails in place to prevent it? If a person with a very good memory reads an article, they only violate…

> Critically the question is, did the developers put reasonable guardrails in place to prevent it? Why? If I steal a bunch of unique works of art and store them in my house for only me to see, am I still committing a crime?

violating copyright is not stealing - it's a government granted monopoly...

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#299
post #293

Earlier quoted context omitted.

Arguing whether it can is not a useful discussion. You can absolutely train a net to memorize and recite text. As these models get more powerful they will memorize more text. The critical thing is how hard is it to make them recite copyrighted works. Critically the question is, did the developers put reasonable guardrails in place to prevent it? If a person with a very good memory reads an article, they only violate…

> Critically the question is, did the developers put reasonable guardrails in place to prevent it? Why? If I steal a bunch of unique works of art and store them in my house for only me to see, am I still committing a crime?

Yes, but policing affairs inside the home have always been impractical at the best of times.

Of course, OpenAI and most other "AI" aren't affairs "inside the home"; they are affairs publicly demonstrated far and wide.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#300

Even if they win against openAI, how would this prevent something like a Chinese or Russian LLM from “stealing” their content and making their own superior LLM that isnt weakened by regulation like the ones in the United States. And I say this as someone that is extremely bothered by how easily mass amounts of open content can just be vacuumed up into a training set with reckless abandon and there isn’t much you can…

The NYT's strongest argument for infringement is that OpenAI is reproducing their content verbatim (and to make matters worse, without attribution). IANAL but it seems super likely to me that this will be found to be infringing sooner or later. Do I really want to use a Chinese word processor that spits unattributed passages from the NYT into the articles I write? Once I publish that to my blog now I'm infringing and…

>IANAL but it seems super likely to me that this will be found to be infringing sooner or later.

It better. Copyright has essentially fucking ceased to exist in the eyes of AI people. Just because you have a shiny new toy doesn't mean the law suddenly stops applying to you. The internet does its best to route around laws and government but the more technologically up to date bureaucracy becomes, the faster it will catch up.

Post reply on HN