Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

441–450 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#441

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

> if NYT goes under a dozen similar outlets can replace them overnight

Not when there’s no money in journalism because the generative AIs immediately steal all content. If nyt goes under no one will be willing to start a news business as everyone will see it’s a money loser.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#442

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

Why using authored NYT articles is “stupid IP battles” and having to pay for the trained model with them is not stupid?

> Why using authored NYT articles is “stupid IP battles”

When an AI uses information from an article it's no difference from me doing it in a blog post. If I'm just summarizing or referencing it, that's fair use, since that's my 'take' on the content.

> having to pay for the trained model with them is not stupid?

Because you can charge for anything you want. I can also charge for my summaries of NYT articles.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#443
post #109

Interesting. I think the appropriation, privatization, and monetization of "all human output" by a single (corporate) entity is at least shameless, probably wrong, and maybe outright disgraceful. But I think OpenAI (or another similar entity) will succeed via the Sackler defense - OpenAI has too many victims for litigation to be feasible for the courts, so the courts must preemptively decide not to bother with compen…

When music people copyright things beats sounds or "style" in music it's even more shameless.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#444
post #99

Earlier quoted context omitted.

If you study copyrighted material for four years at a university and then go on to earn money based on your education, do you owe something to the authors of your text books? I'm not sure how we should treat LLMs with respect to publicly accessible but copyrighted material, but it seems clear to me that "profiting" from copyrighted material isn't a sufficient criteria to cause me to "owe something to the owner".

Do people ever get tired of this argument that relies on anthropomorphizing these AI black boxes? A computer isn't a human, and we already have laws that have a different effect depending on if it's a computer doing it or a human. LLMs are no different, no matter how catchy hyping them up as being == Humans may be.

I didn't anthropomorphize the LLMS. It isn't about laws for the LLM, it is about laws for people would build and operate the LLM.

If you want to assert that groups of people that build and operate LLMs should operate under a different set of laws and regulations than individuals that read books in the library regarding "profit", I'm open to that idea. But that is not at all the same as "anthropomorphizing these AI black boxes".

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#445
post #367

Earlier quoted context omitted.

It doesn't matter what's good for open source ML. It matters what is legal and what makes sense.

The law on this does not currently exist. It is in the process of being created by the courts and legistatures. I personally think that giving copyright holders control over who is legally allowed to view a work that has been made publicly available is a huge step in the wrong direction. One of those reasons is open source, but really that argument applies just as well to making sure that smaller companies have a cha…

It’s disingenuous to frame using data to train a model as a “view,” of that data. The simple cases are the easy ones, if ChatGPT completely rips a NYT article then that’s obviously infringement; however, there’s an argument to be made that every part of the LLM training dataset is, in part, used in every output of that LLM.

I don’t know the solution, but I don’t like the idea that anything I post online that is openly viewable is automatically opted into being part of ML/AI training data, and I imagine that opinion would be amplified if my writing was a product which was being directly threatened by the very same models.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#446
post #383

Earlier quoted context omitted.

> People do the work of turning these happenings into words, and are paid for it. That's what's stolen here. Stolen from whom? Journalists who got reported got paid. The owner is a billionaire. I don't understand your logic. Does NYT pays money to the people/countries etc it uses to as subject to create content(NEWS)? Isn't that stealing then? Also their website TOS didn't prohibit LLMs from using their data.

> Stolen from whom? The owner is a billionaire. > ...owner... > Does NYT pays money to the people/countries etc it uses to as subject to create content(NEWS)? Isn't that stealing then? No, that's why in my reply to "facts like happenings in the world are not copyrightable" I emphasised do the work . Journalism is a job. Happenings do not just fall onto the page. > Also their website TOS didn't prohibit LLMs from usin…

There is no rule of law saying LLMs cannot be trained on WWW data.

New York times made it ridiculously easy for anyone to access their content by putting it in WWW for making money from page impressions. And they started ingesting links of their content to social media, search engines, etc.

And now they are acting surprised someone used the content to train an LLM.

Should have done their job in the first place to prevent it from training LLMs and make it less.

But they didn't because that affects their page impressions and ad views.

Because the more open the content the more money they make everyone click on a link and see the ad.

You can't have it both ways.

If you do gambling by making content so open so you can get more views from ads, you also get to enjoy the consequences and not cry like a baby asking for billions by making stupid decisions in the first place.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#447
post #364

Earlier quoted context omitted.

If the future of humanity rests on access to old NYT articles, we’re fucked. Why can’t OpenAI try to get a license if the NYT archives are so important to them?

They're not. They can skip the entirety of the NYT archives and not much of value will be lost. The issue is with every copycat lawsuit that sues every AI company out of existence. It's a chilling effect on AI development. Old entrenched companies trying to prohibit new ways of learning and sharing information for the sake of their profit.

Why don’t they train their AI on non-copyrighted material? It’s only fair for the copyright owners to want a share of the pie. I’d want one as well for my work.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#448
post #366
post #314

Earlier quoted context omitted.

Playing back large passages of verbatim content sold as your “product” without citation is almost certainly not fair use. Fair use would be saying “The New York Times said X” and then quoting a sentence with attribution. Thats not what OpenAI is being sued for. They’re being sued for passing off substantial bits of NYTimes content as their own IP and then charging for it saying it’s their own IP. This is also related…

At the root, it seems like there's also a gap in copyright with respect to AI around transformative. Is using something, in its entirety, as a tiny bit of a massive data set, in order to produce something novel... infringing? That's a pretty weird question that never existed when copyright was defined.

I think it did come up back in the day sort of, for example with libraries.

More importantly, ever case is unique so what really came up was a set of principles for what defines fair use, which will definitely guide this.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#449

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

Why shouldn't the creators of the training content get anything for their efforts? With some guiderails in place to establish what is fair compensation, Fair Use can remain as-is.

> Why shouldn't the creators of the training content get anything for their efforts?

Well, they didn't charge for it, right? They're retroactively asking for money, but they could have just locked their content behind a strict paywall or had a specific licensing agreement enforceable ahead of time. They could do that going forward, but how is it fair for them to go back and say that?

And the issue isn't "You didn't pay us" it's "This infringes our copyright", which historically the answer has been "no it doesn't".

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#450
post #170

Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.) I don’t necessarily fault OpenAI’s decision to initially train their models without entering into licensing agreements - they probably wouldn’t exist and the generative AI revolution may never have happened if…

It’s likely fair use.

> It's likely fair use.

I agree. You can even listen to the NYT Hard Fork podcast (that I recommend btw https://www.nytimes.com/2023/11/03/podcasts/hard-fork-execut...) where they recently had Harvard copyright law professor Rebecca Tushnet on as a guest.

They asked her about the issue of copyrighted training data. Her response was:

""" Google, for example, with the book project, doesn’t give you the full text and is very careful about not giving you the full text. And the court said that the snippet production, which helps people figure out what the book is about but doesn’t substitute for the book, is a fair use.

So the idea of ingesting large amounts of existing works, and then doing something new with them, I think, is reasonably well established. The question is, of course, whether we think that there’s something uniquely different about LLMs that justifies treating them differently. """

Now for my take: Proving that OpenAI trained on NYT articles is not sufficient IMO. They would need to prove that OpenAI is providing a substitutable good via verbatim copying, which I don't think you can easily prove. It takes a lot of prompt engineering and luck to pull out any verbatim articles. It's well-established that LLMs screw up even well-known facts. It's quite hard to accurately pull out the training data verbatim.

Post reply on HN