Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

381–390 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#381

The arguments about being able to mimic New York Times “style” are weak, but the fact that they got it to emit verbatim NY Times content seems bad for OpenAI: > As outlined in the lawsuit, the Times alleges OpenAI and Microsoft’s large language models (LLMs), which power ChatGPT and Copilot, “can generate output that recites Times content verbatim

I can get a printer to emit verbatim NYT content, and with a lot less effort than getting it out of an LLM. I find this capability of infringement equals infringement argument incredibly weak.

In the EU, countries can (and do) impose levies on printers and scanners because they may be used to copy copyrighted material (https://www.insideglobaltech.com/2013/07/12/eu-member-states...). Similar levies exist for blank CDs, USB sticks, MP3 players etc. In the US, this applies to "blank CDs and personal audio devices, media centers, satellite radio devices, and car audio systems that have recording capabilities." (See https://en.wikipedia.org/wiki/Private_copying_levy)

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#382
post #138
post #109

Interesting. I think the appropriation, privatization, and monetization of "all human output" by a single (corporate) entity is at least shameless, probably wrong, and maybe outright disgraceful. But I think OpenAI (or another similar entity) will succeed via the Sackler defense - OpenAI has too many victims for litigation to be feasible for the courts, so the courts must preemptively decide not to bother with compen…

What do you mean when you say "appropriation and privatization" of "all human output"? The output is still there for anyone else to train on if they want.

> The output is still there for anyone else to train on if they want.

Legal arguments aside, the goldrush era of data scraping is over. Major sources of content like Reddit and Twitter have killed APIs, added defenses and updated EULAs to avoid being pillaged again. More and more sites are moving content behind paywalls.

There's also the small issue of having 10s of millions of VC dollars to rent/buy hundreds of high end GPUs. OpenAI and friends are also trying their hardest to prevent others doing so via 'Skynet' hysteria driven regulatory capture.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#383
post #146

Earlier quoted context omitted.

Yes. Journalism is a job. People do the work of turning these happenings into words, and are paid for it. That's what's stolen here. The value created through doing that work. If it didn't have value, Microsoft would lose nothing by no longer ingesting it.

> People do the work of turning these happenings into words, and are paid for it. That's what's stolen here. Stolen from whom? Journalists who got reported got paid. The owner is a billionaire. I don't understand your logic. Does NYT pays money to the people/countries etc it uses to as subject to create content(NEWS)? Isn't that stealing then? Also their website TOS didn't prohibit LLMs from using their data.

> Stolen from whom? The owner is a billionaire.

> ...owner...

> Does NYT pays money to the people/countries etc it uses to as subject to create content(NEWS)? Isn't that stealing then?

No, that's why in my reply to "facts like happenings in the world are not copyrightable" I emphasised do the work. Journalism is a job. Happenings do not just fall onto the page.

> Also their website TOS didn't prohibit LLMs from using their data.

This is just lazy. We have rule of law. Individuals don't need to write "don't break law X" to be protected by them. And nytimes does in fact have copyright symbols on its pages - not that it needs them.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#384
post #170

Solidly rooting for NYT on this - it’s felt like many creative organizations have been asleep at the wheel while their lunch gets eaten for a second time (the first being at the birth of modern search engines.) I don’t necessarily fault OpenAI’s decision to initially train their models without entering into licensing agreements - they probably wouldn’t exist and the generative AI revolution may never have happened if…

Doesn't this harm open source ML by adding yet another costly barrier to training models?

You can train your own model no problem, but you arguably can’t publish it. So yes, the model can’t be open-sourced, but the training procedure can.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#385
post #370
post #340

Earlier quoted context omitted.

"Why can't AI at least cite its source" each article seen alters the weights a tiny, non-human understandable amount. it doesn't have a source, unless you think of the whole humongous corpus that it is trained on

that just sounds like "we didn't even try to build those systems in that way, and we're all out of ideas, so it basically will never work" which is really just a very, very common story with ai problems, be it sources/citations/licenses/usage tracking/etc., it's all just 'too complex if not impossible to solve', which just seems like a facade for intentionally ignoring those problems for benefit at this point. those…

What makes you think AI researchers (including the big labs like OpenAI and Anthropic) aren't trying to solve these problems?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#386

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

I agree that it’s more important than the NYT, I disagree that it’s the most important development in human history.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#387

Earlier quoted context omitted.

They probably didn’t start with a lawsuit. They started asking for royalties. They probably didn’t get an offer they thought was fair and reasonable so they sued. These media businesses have shareholders and employees to protect. They need to try and survive this technological shift. The internet destroyed their profitability but AI threatens to remove their value proposition.

Sorry, how exactly LLM threatens NYT? Are people supposed to generate news themselves? Or like wait a year or so before NYT articles are consumed by LMMs?

NYT doesn't just publish "news" as in what happened yesterday; they also publish analysis, reviews of books and films, history, biography and so on. That's why people cite NYT articles from decades ago.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#388

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

The way I see it, if the NYT goes under (one of the biggest newspapers in the world), all similar outlets also go under. Major publishers, both of fiction and non-fiction, as well as images, video, and all other creative content, may also go under. Hence, there is no more (reliable) training data.

I'm not sure whether that would even be a net loss, TBH. So much commercial media is crap, maybe it would be better for the profit motive to be removed? On the fiction side, there's plenty of fan-fic and indie productions. On the nonfiction side, many indie creators produce better content these days than the big media outlets do. And there still might be room for premium investigative stories done either by a few consolidated wire outlets (Reuters/APNews) or niche publishers (The Information, 404 Media, etc.).

And then there's all the run-of-the-mill small-town journalism that AI would probably be even better at than human reporters: all the sports stories, the city council meetings, the environmental reviews...

If AI makes commercial content publishing unviable, that might actually cut down on all the SEO spam and make the internet smaller and more local again, which would be a good thing IMO.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#389
post #358
post #331

Earlier quoted context omitted.

For all the leaks on: Secret projects, novelty training algorithms not being published anymore so as to preserve market share, custom hardware, Q* learning, internal politics at companies at the forefront of state of the art LLMs...A thunderous silence is the lack of leaks, on the exact datasets used to train the main commercial LLMs. It is clear OpenAI or Google did not use only Common Crawl. With so many press conf…

Would be fascinated to hear from someone inside on a throwaway, but my nearest experience is that corporate lawyers aren't stupid. If there's legally-murky secret data sauce, it's firewalled from being easily seen in its entirety by anyone not golden-handcuffed to the company. They may be able to train against it. They may be able to peek at portions of it. But no one is downloading-all.

Big corporations and corporate lawyers lose major lawsuits all the time.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#390
post #367

Earlier quoted context omitted.

It doesn't matter what's good for open source ML. It matters what is legal and what makes sense.

The law on this does not currently exist. It is in the process of being created by the courts and legistatures. I personally think that giving copyright holders control over who is legally allowed to view a work that has been made publicly available is a huge step in the wrong direction. One of those reasons is open source, but really that argument applies just as well to making sure that smaller companies have a cha…

Copyright holders already have control over who is legally allowed to view a work that has been made publicly available. It's the right to distribution. You don't waive that right when you make your content free to view on a trial basis to visitors to your site, with the intent of getting subscriptions - however easy your terms are to skirt. NYT has the right to remove any of their content at any time, and to bar others from hosting and profiting on the content.
Post reply on HN