Live data from Hacker News

The New York Times is suing OpenAI and Microsoft for copyright infringement

theverge.com

571–580 of 912 posts

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#571

Earlier quoted context omitted.

> I hope this results in Fair Use being expanded to cover AI training. Couldn't disagree more strongly, and I hope the outcome is the exact opposite. I think we've already started to see the severe negative consequences when the lion's share of the profits get sucked up by very, very few entities (e.g. we used to have tons of local papers and other entities that made money through advertising, now Google and Facebook…

Trying to prohibit this usage of information would not help prevent centralization of power and profit. All it would do is momentarily slow AI progress (which is fine), and allow OpenAI et al to pull the ladder up behind them (which fuels centralization of power and profit). By what mechanism do you think your desired outcome would prevent centralization of profit to the players who are already the largest?

I'm not saying copyright is without problems (e.g. there is no reason I think its protection should be as long as it is), but I think the opposite, where the incentive to create new content (especially in the case of news reporting) is completely killed because someone else gets to vacuum up all the profits, is worse. I mean, existing copyright does protect tons of independent writers, artists, etc. and prevents all of the profits from their output from being "sucked up" by a few entities.

More critically, while fair use decisions are famously a judgement call, I think OpenAI will lose this based on the "effect of the fair use on the potential market" of the original content test. From https://fairuse.stanford.edu/overview/fair-use/four-factors/ :

> Another important fair use factor is whether your use deprives the copyright owner of income or undermines a new or potential market for the copyrighted work. Depriving a copyright owner of income is very likely to trigger a lawsuit. This is true even if you are not competing directly with the original work.

> For example, in one case an artist used a copyrighted photograph without permission as the basis for wood sculptures, copying all elements of the photo. The artist earned several hundred thousand dollars selling the sculptures. When the photographer sued, the artist claimed his sculptures were a fair use because the photographer would never have considered making sculptures. The court disagreed, stating that it did not matter whether the photographer had considered making sculptures; what mattered was that a potential market for sculptures of the photograph existed. (Rogers v. Koons, 960 F.2d 301 (2d Cir. 1992).)

and especially

> “The economic effect of a parody with which we are concerned is not its potential to destroy or diminish the market for the original—any bad review can have that effect—but whether it fulfills the demand for the original.” (Fisher v. Dees, 794 F.2d 432 (9th Cir. 1986).)

The "whether it fulfills the demand of the original" is clearly where NYTimes has the best argument.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#572

Earlier quoted context omitted.

By that logic you should have to pay the copyright holder of every library book you ever read, because you could later produce some content you memorised verbatim.

The rules we have now were made in the context of human brains doing the learning from copyrighted material, not machine learning models. The limitations on what most humans can memorize and reproduce verbatim are extraordinarily different from an LLM. I think it only makes sense to re-explore these topics from a legal point of view given we’ve introduced something totally new.

Human brains are still the main legal agents in play. LLMs are just a computer programs used by humans.

Suppose I research for a book that I'm writing - it doesn't matter whether I type it on a Mac, PC, or typewriter. It doesn't matter if I use the internet or the library. It doesn't matter if I use an AI powered voice-to-text keyboard or an AI assistant.

If I release a book that has a chapter which was blatantly copied from another book, I might be sued under copyright law. That doesn't mean that we should lock me out of the library, or prevent my tools from working there.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#573
post #531
post #480

Earlier quoted context omitted.

LLMs are not databases. There is no "citation" associated with a specific query, any more than you can cite the source of the comment you just made.

That's fine. Solve it a different way. OpenAI doesn't just get to steal work and then say "sorry, not possible" and shrug it off. The NYTimes should be suing.

Clearly, "theft" is an analogy here (since we can't get it to fit exactly), but we can work with it.

You are correct, if I were to steal something, surely I can be made to give it back to you. However, if I haven't actually stolen it, there is nothing for me to return.

By analogy, if OpenAI copied data from the NYT, they should be able to at least provide a reference. But if they don't actually have a proper copy of it, they cannot.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#574

Earlier quoted context omitted.

Oh thank goodness we can rely on charity for our information economy > Also, presumably NYT still has a business model unrelated to whatever OpenAI is doing with [NYT’s] data… That’s exactly the question. They are claiming it is destroying their business, which is pretty much self-evident given all the people in here defending the convenience of OpenAI’s product: they’re getting the fruits of NYTimes’ labor without p…

> Oh thank goodness we can rely on charity for our information economy You seem to be assuming an "information economy" should exist at all. Can you justify that?

We absolutely need an information economy where people can research things and publish what they find without needing some deep pocketed sponsors. Some may do it for money, some may do it for recognition. Once AI absorbs all that information and uses it without attribution these incentives go away. I am sure OpenAI, Microsoft and others will love a world where they have a quasi monopoly on what information goes to the public but I don't think we want that.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#575
post #559

Earlier quoted context omitted.

> They’re being sued for passing off substantial bits of NYTimes content as their own IP and then charging for it saying it’s their own IP. In what sense are they claiming their generated contents as their own IP? https://www.zdnet.com/article/who-owns-the-code-if-chatgpts-... > OpenAI (the company behind ChatGPT) does not claim ownership of generated content. According to their terms of service, "OpenAI hereby assig…

The bits you cite are legally bogus. That would be like me just photocopying a book you wrote and then handing out copies saying we’re assigning different rights to the content. The whole point of the lawsuit is that OpenAI doesn’t own the content and thus they can’t just change the ownership rights per their terms of service. It doesn’t work like that.

Their legalese is careful to include the 'if any' qualifier ("We hereby assign to you all our right, title, and interest, if any, in and to Output.")

In any case, the point is that they made no claim to Output (as opposed to their code, etc) being their IP.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#576
post #99

Earlier quoted context omitted.

If you study copyrighted material for four years at a university and then go on to earn money based on your education, do you owe something to the authors of your text books? I'm not sure how we should treat LLMs with respect to publicly accessible but copyrighted material, but it seems clear to me that "profiting" from copyrighted material isn't a sufficient criteria to cause me to "owe something to the owner".

We don’t, and shouldn’t, give LLMs the same rights as people.

We're not "giving them the same rights as people", we're trying to define the rights of the set of "intelligent" things that can learn (regardless of if their conscious or not). And up until recently, people were the only members of that set.

Now there are (or very, very soon there will be) two members in that set. How do we properly define the rules for members of that set?

If something can learn from reading do ban it from reading copyrighted material, even if it can memorize some of it? Clearly that would be a failure for humans a ban of that form. Should we have that ban for all things that can learn?

There is a reasonable argument that if you want things to learn they have to learn on a wide variety, and on our best works (which are often copyrighted).

And the statements above have no implication of being free of cost (or not), just that I think blocking "learning programs / LLMs" from being able to access, learn from or reproduce copyright text is a net loss for society.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#577

I hope this results in Fair Use being expanded to cover AI training. This is way more important to humanity's future than any single media outlet. If the NYT goes under, a dozen similar outlets can replace them overnight. If we lose AI to stupid IP battles in its infancy, we end up handicapping probably the single most important development in human history just to protect some ancient newspaper. Then another country…

I know utilitarianism is a popular moral theory in hacker circles, but is it really appropriate to dispense with any other notion of justice?

I don’t mean to go off on too deep of a tangent, but if one person’s (or even many people’s) idea of what’s good for humanity is the only consideration for what’s just, it seems clear that the result would be complete chaos.

As it stands, it doesn’t seem to be an “either or” choice. Tech companies have a lot of money. It seems to me that an agreement that’s fundamentally sustainable and fits shared notions of fairness would probably involve some degree of payment. The alternative would be that these resources become inaccessible for LLM training, because they would need to put up a wall or they would go out of business.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#578
post #419

Earlier quoted context omitted.

Just a question, do you remember a source for all the knowledge in your mind, or did you at least try to remember?

No, but I'm a human and treating computers like humans is a huge mistake that we shouldn't make.

Treating computers like humans in this one particular way is very appropriate. It is the only way that LLM can synthesize a worldview when their training data is many thousands of times larger than their number of parameters. Imagine scaling up the total data by another factor of 1million in a few years. There is no current technology to store that info but we can easily train large neural nets that can recreate the essence of it, just like we traditionally trained humans to recall ideas.

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#579

Earlier quoted context omitted.

It's likely not. Search for "the four factors of fair use". While I think OpenAI will have decent arguments for 3 of the factors, they'll get killed on the fourth factor, "the effect of the use on the potential market", which is what this lawsuit is really about. If your "fair use" substantially negatively affects the market for the original source material, which I think is fairly clear in this case, the courts wont…

Fair use is based on a flexible proportionality test so they don't need perfect arguments on all factors. > If your "fair use" substantially negatively affects the market for the original source material, which I think is fairly clear in this case, the courts wont look favorably on that. I think it's fairly clear that it doesn't. No one is going to use ChatGPT to circumvent NYTimes paywalls when archive.ph and the No…

> we're so far in uncharted territory any speculation is useless

I definitely agree with that (at least the "far in uncharted territory bit", but as far as "speculation being useless", we're all pretty much just analyzing/guessing/shooting the shit here, so I'm not sure "usefulness" is the right barometer), which is why I'm looking forward to this case, and I also totally agree the assessment is flexible.

But I don't think your argument that it doesn't negatively affect the market holds water. Courts have held in the past that the market for impact is pretty broadly defined, e.g.

> For example, in one case an artist used a copyrighted photograph without permission as the basis for wood sculptures, copying all elements of the photo. The artist earned several hundred thousand dollars selling the sculptures. When the photographer sued, the artist claimed his sculptures were a fair use because the photographer would never have considered making sculptures. The court disagreed, stating that it did not matter whether the photographer had considered making sculptures; what mattered was that a potential market for sculptures of the photograph existed. (Rogers v. Koons, 960 F.2d 301 (2d Cir. 1992).)

From https://fairuse.stanford.edu/overview/fair-use/four-factors/

Re: The New York Times is suing OpenAI and Microsoft for copyright infringement

#580
post #340

Earlier quoted context omitted.

"Why can't AI at least cite its source" each article seen alters the weights a tiny, non-human understandable amount. it doesn't have a source, unless you think of the whole humongous corpus that it is trained on

So why my employer implementation version of azure chatgpt on our document systems can successfully cite its sourced documents?

My understanding is that this lawsuit is about the training corpus. This is on the level of asking it to cite its sources for a/an/the.
Post reply on HN