Live data from Hacker News

New York Times considers legal action against OpenAI as copyright tensions swirl

npr.org

261–270 of 383 posts

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#261

Earlier quoted context omitted.

Fair use also covers generating derived works, and arguably the AI is a derived work.

That would be hard to argue. You can't point to a sequence of words that's in the original that was copied into the OpenAI output.

That's not how derivative works are defined. For example, translations to other languages or extensive summarizations (condensations) count as derivative works even though no words are directly copied.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#262
post #252

Earlier quoted context omitted.

> AI consuming copyrighted data and producing an output has to be considered a derivative work (or indeed, the model itself will be considered a derivative work) or IP protections are effectively broken. It’s not derivative work though. First, a human didn’t create it, so copyright protections don’t exist on its output. Machines don’t enjoy copyright protections, people do. It’s mechanically copying and reproducing p…

Seems unfair as we converge on AGI. If I memorize the lyrics to a song, is that a copyright violation? The lyrics are encoded in the arrangement of my neurons, after all.

I don't know why people use these analogies. No person can memorize terabytes worth of lyrics.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#263

Honestly, I think generative AI losing a massive copyright showdown is inevitable at this stage. It's extremely easy to get the latest generation of AIs to produce outputs that in many fields sans-AI would be trivially considered as IP infringement. While there are many interesting reasonable legal & technical arguments that it's not, the result completely undermines copyright protections regardless. If that's accept…

> AI consuming copyrighted data and producing an output has to be considered a derivative work (or indeed, the model itself will be considered a derivative work) or IP protections are effectively broken. It’s not derivative work though. First, a human didn’t create it, so copyright protections don’t exist on its output. Machines don’t enjoy copyright protections, people do. It’s mechanically copying and reproducing p…

I think attempts to tease out this distinction is going to make these laws unwieldy to use in practice

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#264
>>A top concern for The Times is that ChatGPT is, in a sense, becoming a direct competitor with the paper by creating text that answers questions based on the original reporting and writing of the paper's staff.

This seems to me to be completely standard in the newspaper industry. Many times every week, I see stories in the form "The [Major_News_Outlet] reports that [Event_X occurred] or [their investigation revealed Y] and here are the details [...].

Copyright protects the expression of an idea, not the idea itself. If you write a history of Issac Newton or the invention of semiconductors, I cannot copy that wholesale and sell it as mine, but nothing prevents me writing my own version, even using the same facts and citing your work.

I'm quite sure that I could provide a service where a bunch of workers read NYT articles and write brief summaries. I'm not sure they would even need citations, as long as we don't copy chunks wholesale.

If OpenAI is simply parroting the words of the NYT articles without Fair Use constraints (short blurbs), it seems they have a problem. If they are fully re-writing them into short non-copying summaries, it seems the NYT has a problem.

It'll be interesting to see how the courts sort this out.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#265

IANAL, but copyright protections are pretty much tied to content and format and not to the idea itself, with the intent of preventing (or putting a price on) the copying of original works. The Times will have a very hard time proving that their content is being re-marketed by OpenAI. Having a competing product based on your ideas. Compare: "Steve Jobs [was] a tyrant": https://www.nytimes.com/2011/10/07/technology/ste…

The NYT argument is going to be that they put up a site, own the copyright for their content and make that content available for either a human to read it for themselves, or software to index for something commonly understood as a search engine. Those terms do not entitle the training of LLMs for commercial use. Therefore, cease and desist. Oh and destroy anything that was created by violating the terms of our licens…

Terms of Use are a thing, and if the Times can prove that OpenAI infringed their web terms by scraping, they may have a case... but terms of use probably won't monetize well or give them enough leverage to prevent OpenAI from using their data anyway and may end-up distracting from the main copyright suit.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#266

Earlier quoted context omitted.

> But OpenAI is neither copying nor deriving a work. They absolutely are copying in the course of training the model, and they are doing something that often looks a lot like copying when producing output with the model. > Style (which can be described as a probability model) is not copyrightable. Style is not all that can be described in an LLM's "probability model", otherwise LLM models would never be able to repro…

> They absolutely are copying in the course of training the model Only in the sense that a Cisco router is copying in the course of sending me the article, which we've all agreed doesn't count as infringement. The bigger problem is that the plaintiff has to show it is more likely than not that that sequence of words came from their text and not some other source , which is going to be obscenely difficult. > And ChatG…

> Only in the sense that a Cisco router is copying in the course of sending me the article, which we've all agreed doesn't count as infringement.

But it does count as infringement, if its not explicltly or implicitly licensed (because it is necessary to a use that is licensed or necessary to a use that does not itself require a license but is the normal use for for which a licensed copy, which you have, is sold and used) and it is not itself Fair Use (usually, because it is necessary in the course of a use which is itself Fair Use.)

> The bigger problem is that the plaintiff has to show it is more likely than not that that sequence of words came from their text and not some other source, which is going to be obscenely difficult.

The plaintiff has to (1) show facts from which a judge concludes that a reasonable jury might conclude that, and (2) get the jury to conclude that.

This can be difficult, and it might not be trivial in this case, but in practice the combination of opportunity and non-trivial similarity tends to put more weight on the defense to show a convincing alternative explanation. Outside of conclusions that a judge views as conpletely unreasonable based on the facts, the civil burden of proof boils down to what a jury feels is more likely, not some.

> Granted, and then the copyright holder could sue for infringement at that point. Exactly whom he should sue is a more difficult question.

Well, who else might be liable depends on the specific circumstances, but with OpenAI controlling the whole course between the first copy of the copyrighted work and the infringing end product, and doing it all for commercial gain, that OpenAI wpuld be on the hook, civilly and potentially criminally for any infringement, is clear.

> Look, I get it: you want copyright to allow authors to say "you can't do that with my work"; I'm sympathetic, but it just doesn't give authors that power.

I want copyright to be both more limited in its exclusive rights and either shorter or costlier to the copyright holder than it is. But what I want is not what the law is.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#267

Honestly, I think generative AI losing a massive copyright showdown is inevitable at this stage. It's extremely easy to get the latest generation of AIs to produce outputs that in many fields sans-AI would be trivially considered as IP infringement. While there are many interesting reasonable legal & technical arguments that it's not, the result completely undermines copyright protections regardless. If that's accept…

> AI consuming copyrighted data and producing an output has to be considered a derivative work (or indeed, the model itself will be considered a derivative work) or IP protections are effectively broken. It’s not derivative work though. First, a human didn’t create it, so copyright protections don’t exist on its output. Machines don’t enjoy copyright protections, people do. It’s mechanically copying and reproducing p…

> First, a human didn’t create it, so copyright protections don’t exist on its output. Machines don’t enjoy copyright protections, people do.

This is not correct. AI models are tools that humans use.

This is like saying "it was typed on a computer therefore it doesn't enjoy copyright protections"

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#268
post #252

Earlier quoted context omitted.

Seems unfair as we converge on AGI. If I memorize the lyrics to a song, is that a copyright violation? The lyrics are encoded in the arrangement of my neurons, after all.

I don't know why people use these analogies. No person can memorize terabytes worth of lyrics.

What if someone could? What if rainman could

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#270
post #255

Earlier quoted context omitted.

Eh, I could definitely see the artists losing. The most obvious scenario that comes to mind for me is, imagine an independent artists launching their (book/film/album/etc) and the same day someone with more resources and experience takes the work and markets it better than the OG author ever could on their own.

It would be great if such "creative" works were simply impossible to monetize. I already don't pay for these and try to find unknown artists / writers who have a job and do stuff simply because they enjoy doing it.

That’s… awful.
Post reply on HN