Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

401–410 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#401

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

Which dozen outlets can replace the New York Times overnight? I will stipulate that the NYT isn’t worthy of historic preservation if it’s become obsolete — but which dozen outlets can replace it? Wouldn’t those dozen outlets suffer the same harms of producing original content, costing time and talent, and while having a significant portion of the benefit accruing to downstream AI companies? If most of the benefit of…

Just pick the top 12 articles/publishers out of a month of Google News, doesn't really matter. Most readers probably can't tell them apart anyway.

Yes, all those outlets will suffer the same harms. They have been for decades. That's why there's so few remaining. Most are consolidated and produce worthless drivel now. Their business model doesn't really work in the modern era.

Thankfully, people have and will continue to produce content even if much of it gets stolen -- as has happened for decades, if not millennia, before AI.

If anything what we need is a better way to fund human creative endeavors not dependent on pay-per-view. That's got nothing to do with AI; AI just speeds up a process of decay that has been going on forever.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#402
post #340

Earlier quoted context omitted.

"Why can't AI at least cite its source" each article seen alters the weights a tiny, non-human understandable amount. it doesn't have a source, unless you think of the whole humongous corpus that it is trained on

So why my employer implementation version of azure chatgpt on our document systems can successfully cite its sourced documents?

Because the model proper wasn’t trained on those documents, it’s just RAG being employed with the documents as external sources. It’s a fundamentally different setup.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#403
post #123

Google can look up into their index and can remove whatever they want to, within minutes. But how that can be possible for an LLM? That is, "decontaminate" the model from certain parts of the corups? I can only think of excluding the data set from the training and then retrain? As a side note, I think LLM frenzy would be dead in few years, 10 years time frame at max. The rent seeking on these LLMs as of today would n…

But, wasn’t the reason proprietary unixes died out at major work horses because of a nearly feature comparable free alternative (Linux)? Extending the analogy, LLMs won’t die out, just proprietary ones. (Which is where I think this tech will actually go anyway.)

[deleted]

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#404
post #314

Earlier quoted context omitted.

It’s likely fair use.

Playing back large passages of verbatim content sold as your “product” without citation is almost certainly not fair use. Fair use would be saying “The New York Times said X” and then quoting a sentence with attribution. Thats not what OpenAI is being sued for. They’re being sued for passing off substantial bits of NYTimes content as their own IP and then charging for it saying it’s their own IP. This is also related…

> They’re being sued for passing off substantial bits of NYTimes content as their own IP and then charging for it saying it’s their own IP.

In what sense are they claiming their generated contents as their own IP?

https://www.zdnet.com/article/who-owns-the-code-if-chatgpts-...

> OpenAI (the company behind ChatGPT) does not claim ownership of generated content. According to their terms of service, "OpenAI hereby assigns to you all its right, title and interest in and to Output."

https://openai.com/policies/terms-of-use

> Ownership of Content. As between you and OpenAI, and to the extent permitted by applicable law, you (a) retain your ownership rights in Input and (b) own the Output. We hereby assign to you all our right, title, and interest, if any, in and to Output.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#405
post #123

Google can look up into their index and can remove whatever they want to, within minutes. But how that can be possible for an LLM? That is, "decontaminate" the model from certain parts of the corups? I can only think of excluding the data set from the training and then retrain? As a side note, I think LLM frenzy would be dead in few years, 10 years time frame at max. The rent seeking on these LLMs as of today would n…

Windows and MacOS (and their closed source derivatives) are probably at least as large as Linux, even including all the servers Linux is deployed on. Proprietary UNIX did not "die out"; Apple sells about a quarter million of them every year.

The majority of the world's computing systems runs on closed source software. Believing the opposite is bubble-thinking. Its not just Windows/MacOS. Most Android distros are not actually open source. Power control systems. Traffic control systems. Networking hardware. Even the underlying operating systems which power the VMs you run on AWS that are technically open source. The billions of little computers that form together to make the modern world work; they're mostly closed source.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#406
post #327

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

Why can't AI at least cite its source? This feels like a broader problem, nothing specific to the NYTimes. Long term, if no one is given credit for their research, either the creators will start to wall off their content or not create at all. Both options would be sad. A humane attribution comment from the AI could go a long way - "I think I read something about this in the NYTimes on January 3rd, 2021." It appears t…

I think the gap between attributable knowledge and absorbed knowledge is pretty difficult to bridge. For news stuff, if I read the same general story from NYT and LA Times and WaPo then I'll start to get confused about which bit I got from which publication. In some ways, being able to verbatim quote long passages is a failure to generalize that should be fixed rather than reinforced.

Though the other way to do it is to clearly document the training data as a whole, even if you can't cite a specific entry in it for a particular bit of generated output. It should get useless quickly though as you'd eventually have one big citation -- "The Internet"

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#407

Earlier quoted context omitted.

It is neutral though. That’s the whole point. You have to twist its arm with great intention to recreate specific things. Sufficient intention that it’s really on you at that point.

It’s not neutral if all the content is in the model, regardless of whether you had to twist its arm or not. What does that even mean with a piece of software? A printer is neutral because you have to send it all the data to print out a copy of copyrighted content. It doesn’t contain it inherently.

The fact that the NYT lawyers used a carefully written prompt kind of nullifies this argument. It's not like they stumbled on it on accident, they looked for it and their prompt isn't neutral either.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#408
post #314

Earlier quoted context omitted.

It’s likely fair use.

Playing back large passages of verbatim content sold as your “product” without citation is almost certainly not fair use. Fair use would be saying “The New York Times said X” and then quoting a sentence with attribution. Thats not what OpenAI is being sued for. They’re being sued for passing off substantial bits of NYTimes content as their own IP and then charging for it saying it’s their own IP. This is also related…

I would note that in the examples the NYT cites, the prompts explicitly ask for the reproduction of content.

I think it makes sense to hold model makers responsible when their tools make infringement too easy to do or possible to do accidentally. However that is a far cry from requiring a little longer license to do the trainint in the first place.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#409
post #321

Earlier quoted context omitted.

Except that every stackoverflow post is explicitly creative commons: https://stackoverflow.com/help/licensing

So I suppose it would be the like saying that if you used Stack Overflow to find answers, all of the work you created using information from it would have to be explicitly under the Creative Commons license. You wouldn't even be able to work for companies who aren't using that license if some of your knowledge comes from what you learned on Stack Overflow. Used Stack Overflow to learn anything about programming? You'…

If your code contains verbatim copy-paste of entire blocks of non-trivial code lifted from those videos/books/newsletters with commercial licenses, then yes you would be liable for some licensing fees, at minimum.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#410
post #396
post #327

Earlier quoted context omitted.

Why can't AI at least cite its source? This feels like a broader problem, nothing specific to the NYTimes. Long term, if no one is given credit for their research, either the creators will start to wall off their content or not create at all. Both options would be sad. A humane attribution comment from the AI could go a long way - "I think I read something about this in the NYTimes on January 3rd, 2021." It appears t…

If you're going to consider training ai as fair use, you'll have all kinds of different people with different skill levels training ais that work in different ways on the corpus. Not all of them will have the capability to cite a source, and plenty of them won't have it make sense to cite a source. Eg. Suppose I train a regression that guesses how many words will be in a book. Which book do I cite when I do an infere…

Any citation would be a good start.

For complex subjects, I'm sure the citation page would be large, and a count would be displayed demonstrating the depth of the subject[3].

This is how Google did it with search results in the early days[1]. Most probable to least probable, in terms of the relevancy of the page. With a count of all possible results [2].

The same attempt should be made for citations.

Post reply on HN