Live data from Hacker News

New York Times considers legal action against OpenAI as copyright tensions swirl

npr.org

221–230 of 383 posts

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#221
post #213

IANAL, but copyright protections are pretty much tied to content and format and not to the idea itself, with the intent of preventing (or putting a price on) the copying of original works. The Times will have a very hard time proving that their content is being re-marketed by OpenAI. Having a competing product based on your ideas. Compare: "Steve Jobs [was] a tyrant": https://www.nytimes.com/2011/10/07/technology/ste…

> A top concern for The Times is that ChatGPT is, in a sense, becoming a direct competitor with the paper by creating text that answers questions based on the original reporting and writing of the paper's staff Sounds to me like they are trying to claim copyright over facts rather than the specific expression. That’s just not how copyright works at the moment. The framing of openAIs recent changes is telling too. Ope…

>which the press is now framing as “trying to hide the use of copyrighted data”

Yea, now I can't read the paper and talk about it to other people it seems.

The Right to Read was a prophecy I guess?

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#222

IANAL, but copyright protections are pretty much tied to content and format and not to the idea itself, with the intent of preventing (or putting a price on) the copying of original works. The Times will have a very hard time proving that their content is being re-marketed by OpenAI. Having a competing product based on your ideas. Compare: "Steve Jobs [was] a tyrant": https://www.nytimes.com/2011/10/07/technology/ste…

I think this is going to be a test of the fair use doctrine. https://www.copyright.gov/fair-use/ Now, there's this idea that "news" is just factual and therefore falls under "fair use". However, that's only part of what section 107 says. Fair use very much is still conditional, as there are 4 factors to be considered: (a) Purpose and character of the use, including whether the use is of a commercial nature or is for…

Unfortunately we'll have 9 justices who know little about copyright law and nothing about tech to tell us what the law really is.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#223
post #184

Earlier quoted context omitted.

IMHO no because LLM's are not legal persons and most likely will never be, so they can't acquire or operate under laws, privileges and agreements simply because they exists.

Yeah I agree that legal personhood for LLMs at this point is far-fetched. This would be a separate argument, though, from the notion I responded to above that the difference in the processes of a human mind and LLM are the reason why "learning" from copyrighted material is a violation of copyright in one case and not the other.

(IANAL) I'm not sure if I understand correctly, but I don't think so. Since LLM is not a person in the legal term it really doesn't matter what the difference is. There may be virtually no difference but I imagine the discussion would still be academic. For example: animals aren't granted rights just because they are in some instances similar or in other instances even identical to humans. Primates aren't allowed to walk everywhere humans can just because they posses the ability to walk on two legs.

In my view the biggest issue to raise is the effect of the use upon the potential market for or value of the copyrighted work (https://en.wikipedia.org/wiki/Fair_use). The most spectacular example of damage would be stackoverflow, thought stackoverflow content is not copyrighted. I think there is little doubt that LLM's drive attention from original sources. That might be deemed damaging, especially in the long run.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#224
post #184

Earlier quoted context omitted.

So if we add a bunch of complex processes to an LLM, in order to produce a better analog of a human in terms of degree of complexity if not actual function, does that have some bearing on this copyright question? It doesn’t seem clear to me that it does. Is the argument that sufficient complexity in how an “intelligence” processes this copyrighted data leads to the output being transformative vs not transformative in…

IMHO no because LLM's are not legal persons and most likely will never be, so they can't acquire or operate under laws, privileges and agreements simply because they exists.

But a business can (at least in the US)?

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#225

Honestly, I think generative AI losing a massive copyright showdown is inevitable at this stage. It's extremely easy to get the latest generation of AIs to produce outputs that in many fields sans-AI would be trivially considered as IP infringement. While there are many interesting reasonable legal & technical arguments that it's not, the result completely undermines copyright protections regardless. If that's accept…

Ruling in favor of copyright will call into question search engines and the like as well. Do you think Bing or Google are going to negotiate copying rights with the world's websites? LLMs are proving that intellectual property has a bunch of holes in it. It's been unstable ground to defend since day one. Upon what principle should we believe that one can own an idea and all performances or derivatives of it? Patents…

I think there is a major qualitative difference between generative AI and search engines.

Search engines index the web and point you at other people's work, along the way showing perhaps too much of that content (thus "stealing" users from the target webpage). But they don't reshuffle existing content into something apparently new and original.

The "malicious" case for generative AI is that it sucks in copyrighted work (vs. indexing it), rehashes and produces something that is supposedly original, but really a sophisticated rehash of copyrighted work.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#226

Earlier quoted context omitted.

If it’s on the open internet then why should they have to do that? How is openai training on articles fundamentally different from the wayback machine storing them? They’re just getting stored in a different form.

Copyright law includes many exceptions explicitly for libraries.

Copywrite law doesn't ban people from reading the source altogether.

That is what is being proposed, ban AI from being allowed to 'read' the content.

The real argument is how does a human brain aggregate knowledge and then profit from it, and is it really that different from an AI model aggregating knowledge.

They both read in data, perform calculations on the data, and spit out something.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#227
post #224
post #184

Earlier quoted context omitted.

IMHO no because LLM's are not legal persons and most likely will never be, so they can't acquire or operate under laws, privileges and agreements simply because they exists.

But a business can (at least in the US)?

But by business we mean an organization of people.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#228
post #221
post #213

Earlier quoted context omitted.

> A top concern for The Times is that ChatGPT is, in a sense, becoming a direct competitor with the paper by creating text that answers questions based on the original reporting and writing of the paper's staff Sounds to me like they are trying to claim copyright over facts rather than the specific expression. That’s just not how copyright works at the moment. The framing of openAIs recent changes is telling too. Ope…

>which the press is now framing as “trying to hide the use of copyrighted data” Yea, now I can't read the paper and talk about it to other people it seems. The Right to Read was a prophecy I guess?

An individual or group of individuals doing this and sharing their views/summary vs. a profit-oriented program funded by major technology companies scraping this information and spitting it back out algorithmically does seem different to me. Yes, perhaps both things are on the same "sliding scale", but I do not view them as fundamentally equivalent actions.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#229
post #221

Earlier quoted context omitted.

>which the press is now framing as “trying to hide the use of copyrighted data” Yea, now I can't read the paper and talk about it to other people it seems. The Right to Read was a prophecy I guess?

An individual or group of individuals doing this and sharing their views/summary vs. a profit-oriented program funded by major technology companies scraping this information and spitting it back out algorithmically does seem different to me. Yes, perhaps both things are on the same "sliding scale", but I do not view them as fundamentally equivalent actions.

>a profit-oriented program

The particular problem here is this program isn't magic, it just requires a lot of electricity and hardware to train at the moment. If at some point in the future this hardware becomes cheap then now suddenly OSS LLMs would be under the same set of rules that we're applying to major technology companies.

But mark my words, the large copyright holding groups don't give any shits other than how much IP they can scrape up and demand money for, for the next few human lifetimes.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#230

IANAL, but copyright protections are pretty much tied to content and format and not to the idea itself, with the intent of preventing (or putting a price on) the copying of original works. The Times will have a very hard time proving that their content is being re-marketed by OpenAI. Having a competing product based on your ideas. Compare: "Steve Jobs [was] a tyrant": https://www.nytimes.com/2011/10/07/technology/ste…

Can't wait for the supreme Court ruling that says AI is just using data, and data is free.
Post reply on HN