Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

361–370 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#361

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

Which dozen outlets can replace the New York Times overnight? I will stipulate that the NYT isn’t worthy of historic preservation if it’s become obsolete — but which dozen outlets can replace it?

Wouldn’t those dozen outlets suffer the same harms of producing original content, costing time and talent, and while having a significant portion of the benefit accruing to downstream AI companies?

If most of the benefit of producing original content accrues to the AI firms, won’t original content stop being produced?

If original content stops being produced, how will AI models get better in the future?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#362
post #347

Earlier quoted context omitted.

It doesn't matter what's good for open source ML. It matters what is legal and what makes sense.

Slavery was legal...

Still is in many countries with excellent diplomatic relations with the Western World:

https://www.cfr.org/backgrounder/what-kafala-system

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#363
post #357
post #340

Earlier quoted context omitted.

"Why can't AI at least cite its source" each article seen alters the weights a tiny, non-human understandable amount. it doesn't have a source, unless you think of the whole humongous corpus that it is trained on

We're trying to solve AGI but can't solve sources/citations?

[deleted]

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#364

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

If the future of humanity rests on access to old NYT articles, we’re fucked. Why can’t OpenAI try to get a license if the NYT archives are so important to them?

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#365

Earlier quoted context omitted.

I can get a printer to emit verbatim NYT content, and with a lot less effort than getting it out of an LLM. I find this capability of infringement equals infringement argument incredibly weak.

Well imagine you sell a printer with internal memory loaded with NYT content

[dead]

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#366
post #314

Earlier quoted context omitted.

It’s likely fair use.

Playing back large passages of verbatim content sold as your “product” without citation is almost certainly not fair use. Fair use would be saying “The New York Times said X” and then quoting a sentence with attribution. Thats not what OpenAI is being sued for. They’re being sued for passing off substantial bits of NYTimes content as their own IP and then charging for it saying it’s their own IP. This is also related…

At the root, it seems like there's also a gap in copyright with respect to AI around transformative.

Is using something, in its entirety, as a tiny bit of a massive data set, in order to produce something novel... infringing?

That's a pretty weird question that never existed when copyright was defined.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#367

Earlier quoted context omitted.

Doesn't this harm open source ML by adding yet another costly barrier to training models?

It doesn't matter what's good for open source ML. It matters what is legal and what makes sense.

The law on this does not currently exist. It is in the process of being created by the courts and legistatures.

I personally think that giving copyright holders control over who is legally allowed to view a work that has been made publicly available is a huge step in the wrong direction. One of those reasons is open source, but really that argument applies just as well to making sure that smaller companies have a chance of competing.

I think it makes much more sense to go after the infringing uses of models rather than putting in another barrier that will further advantage the big players in this space.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#368

Earlier quoted context omitted.

Establishing a legal route to train LLMs on copywriten content could certainly have a chilling affect on the progress of science and useful arts... Why would someone devote their life to their studies or craft when they know that an LLM will hoover it up and start plagiarizing it immediately?

The vast majority of quality art and is produced by people who do it because they want to create art, not for money, and most artists earn little.

Artists getting paid little is a function of how easy it is to capture the value that artists create, and how willing other people are to do that capturing - not how valuable their work is. Artists definitely want to get paid and do not want to live on scraps for "passion's" sake. This is the entire argument for copyright existing in the first place, even though current copyright law has flaws.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#369
post #327

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

Why can't AI at least cite its source? This feels like a broader problem, nothing specific to the NYTimes. Long term, if no one is given credit for their research, either the creators will start to wall off their content or not create at all. Both options would be sad. A humane attribution comment from the AI could go a long way - "I think I read something about this in the NYTimes on January 3rd, 2021." It appears t…

There's a few levels to this...

Would it be more rigorous for AI to cite its sources? Sure, but the same could be said for humans too. Wikipedia editors, scholars, and scientists all still struggle with proper citations. NYT itself has been caught plagiarizing[1].

But that doesn't really solve the underlying issue here: That our copyright laws and monetization models predate the Internet and the ease of sharing/paywall bypass/piracy. The models that made sense when publishing was difficult and required capital-intensive presses don't necessarily make sense in the copy and paste world of today. Whether it's journalists or academics fighting over scraps just for first authorship (while some random web dev makes 3x more money on ad tracking), it's just not a long-term sustainable way to run an information economy.

I'd also argue that attribution isn't really that important to most people to begin with. Stuff, real and fake, gets shared on social media all the time with limited fact-checking (for better or worse). In general, people don't speak in a rigorous scholarly way. And people are often wrong, with faulty memories, or even incentivized falsehoods. Our primate brains aren't constantly in fact-checking mode and we respond better to emotional, plot-driven narratives than cold statistics. There are some intellectuals who really care deeply about attributions, but most humans won't.

Taken the above into consideration:

1) Useful AI does not necessarily require attribution

2) AI piracy is just a continuation of decades of digital piracy, and the solutions that didn't work in the 1990s and 2000s still won't work against AI

3) We need some better way to fund human creativity, especially as it gets more and more commoditized

4) This is going to happen with or without us. Cat's outta the bag.

I don't think using old IP law to hold us back is really going to solve anything in the long term. Yes, it'd be classy of OpenAI to pay everyone it sourced from, but long term that doesn't matter. Creativity has always been shared and copied and imitated and stolen, the only question is whether the creators get compensated (or even enriched) in the meantime. Sometimes yes, sometimes no, but it happens regardless. There'll always be noncommercial posts by the billions of people who don't care if AI, or a search engine, or Twitter, or whoever, profits off them.

If we get anywhere remotely close to AGI, a lot of this won't matter. Our entire economic and legal systems will have to be redone. Maybe we can finally get rid of the capitalist and lawyer classes. Or they'll probably just further enslave the rest of us with the help of their robo-bros, giving AI more rights than poor people.

But either way, this is way bigger than the economics of 19th-century newspapers...

[1] https://en.wikipedia.org/wiki/Jayson_Blair#Plagiarism_and_fa...

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#370
post #340
post #327

Earlier quoted context omitted.

Why can't AI at least cite its source? This feels like a broader problem, nothing specific to the NYTimes. Long term, if no one is given credit for their research, either the creators will start to wall off their content or not create at all. Both options would be sad. A humane attribution comment from the AI could go a long way - "I think I read something about this in the NYTimes on January 3rd, 2021." It appears t…

"Why can't AI at least cite its source" each article seen alters the weights a tiny, non-human understandable amount. it doesn't have a source, unless you think of the whole humongous corpus that it is trained on

that just sounds like "we didn't even try to build those systems in that way, and we're all out of ideas, so it basically will never work"

which is really just a very, very common story with ai problems, be it sources/citations/licenses/usage tracking/etc., it's all just 'too complex if not impossible to solve', which just seems like a facade for intentionally ignoring those problems for benefit at this point. those problems definitely exist, why not try to solve them? because well...actually trying to solve them would entail having to use data properly and pay creators, and that'd just cut into bottom line. the point is free data use without having to pay, so why would they try to ruin that for themselves?

Post reply on HN