Live data from Hacker News

New York Times considers legal action against OpenAI as copyright tensions swirl

npr.org

181–190 of 383 posts

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#181

IANAL, but copyright protections are pretty much tied to content and format and not to the idea itself, with the intent of preventing (or putting a price on) the copying of original works. The Times will have a very hard time proving that their content is being re-marketed by OpenAI. Having a competing product based on your ideas. Compare: "Steve Jobs [was] a tyrant": https://www.nytimes.com/2011/10/07/technology/ste…

Discovery is how they will show that it’s their news; and if it gets that far we might finally learn what data they trained it on and how much.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#182

Honestly, I think generative AI losing a massive copyright showdown is inevitable at this stage. It's extremely easy to get the latest generation of AIs to produce outputs that in many fields sans-AI would be trivially considered as IP infringement. While there are many interesting reasonable legal & technical arguments that it's not, the result completely undermines copyright protections regardless. If that's accept…

Or on the converse: if those industries are unviable without copyright protection, they could go away entirely. This is a plausible path to "drop copyright entirely", just like encryption was dropped as an export-controlled technology in the late 90s. (remember the 40-bit "international" SSL?)

OpenAI etc. have huge amounts of money behind them, they very well have a fighting chance in court to defend their usage of scraping the internet.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#183
post #163

Earlier quoted context omitted.

The problem with the "a lossy mathematical translation of its inputs is exactly like a person learns" arguments, even if courts don't find them ludicrous, is that people absolutely can and are found guilty of trademark violations when they read thousands of pages of the LOTR and then write a fantasy novel full of Tolkein's character names for profit.

A fantasy novel with Tolkien's characters' names is an evident copyright violation regardless of how it was generated. That's not what's happening here.

No, but the point is that OpenAI expects to be held harmless if someone uses its tools to violate trademarks [unlike if a human employee had been paid to read Tolkein and write a story about Gandalf and hobbits] because its not an OpenAI employee consciously doing it, but just a model that translates its given inputs

If GPT is blameless doing some things because it's a deterministic model, not an agent, then the "it would be okay if a person was taught like this" defence doesn't apply in other areas

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#184

Earlier quoted context omitted.

Computers aren't humans and LLMs aren't human brains. We have no way to reconstruct memories from a preserved brain (yet). The exact ways in which humans form memories and store information isn't even known yet; we're still drilling into the specifics from higher-level concepts. Modeling the human brain like nodes with weights ignores a lot of biological processes. Blood/oxygen flow, hormones, neurotransmitter decay,…

So if we add a bunch of complex processes to an LLM, in order to produce a better analog of a human in terms of degree of complexity if not actual function, does that have some bearing on this copyright question? It doesn’t seem clear to me that it does. Is the argument that sufficient complexity in how an “intelligence” processes this copyrighted data leads to the output being transformative vs not transformative in…

IMHO no because LLM's are not legal persons and most likely will never be, so they can't acquire or operate under laws, privileges and agreements simply because they exists.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#186

Earlier quoted context omitted.

> Don't humans operate similarly? I'm going to bypass this question a bit and say, who cares? Why do we need to treat these things the same way we treat humans? Why can we not say that it's okay if a human does it, and not okay if it's a computer? There's nothing that requires us to establish 'fair' as treating them the same as people.

> Why do we need to treat these things the same way we treat humans? Because it would be absurd for it to be legal to do X, but illegal to do so with an efficient tool. Especially when the activity X in question is "learning".

> Especially when the activity X in question is "learning".

No such thing, "learning" in a vacuum describes nothing here. Might as well ask why I am allowed to make noise, e.g. speak, but when I install 5000 watt speakers on every square meter of the planet suddenly it's a problem, and roll my eyes at the inconsistency of not being allowed to "do X more efficiently with a tool".

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#187
post #171
post #142

Earlier quoted context omitted.

> A ruling in favor of copyright would force OpenAI to shut down - but given their impressive demo of the tech Who cares? If they want data, they can pay for it.

Or at least ask before scrapping/reading it.

If it’s on the open internet then why should they have to do that? How is openai training on articles fundamentally different from the wayback machine storing them? They’re just getting stored in a different form.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#188

IANAL, but copyright protections are pretty much tied to content and format and not to the idea itself, with the intent of preventing (or putting a price on) the copying of original works. The Times will have a very hard time proving that their content is being re-marketed by OpenAI. Having a competing product based on your ideas. Compare: "Steve Jobs [was] a tyrant": https://www.nytimes.com/2011/10/07/technology/ste…

they had to make a copy of the original to get it into their system in the first place!

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#189

IANAL, but copyright protections are pretty much tied to content and format and not to the idea itself, with the intent of preventing (or putting a price on) the copying of original works. The Times will have a very hard time proving that their content is being re-marketed by OpenAI. Having a competing product based on your ideas. Compare: "Steve Jobs [was] a tyrant": https://www.nytimes.com/2011/10/07/technology/ste…

they had to make a copy of the original to get it into their system in the first place!

In order to render that page, it probably was copied dozens of times all over my RAM. Do I owe NYT money now?

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#190

Honestly, I think generative AI losing a massive copyright showdown is inevitable at this stage. It's extremely easy to get the latest generation of AIs to produce outputs that in many fields sans-AI would be trivially considered as IP infringement. While there are many interesting reasonable legal & technical arguments that it's not, the result completely undermines copyright protections regardless. If that's accept…

Ruling in favor of copyright will call into question search engines and the like as well.

Do you think Bing or Google are going to negotiate copying rights with the world's websites?

LLMs are proving that intellectual property has a bunch of holes in it. It's been unstable ground to defend since day one. Upon what principle should we believe that one can own an idea and all performances or derivatives of it? Patents and trademarks haven't really helped as much as they were expected to. Only recently did works as far back as 1920 enter the public domain. Additionally, some LLMs gobble up source code with mixed licensing structures. How is that to be handled?

Patents are about monopolies to produce something to hit a market, with the trade-off of showing everyone how it's made.

Trademarks just allow you to defend your name(s).

Copyright prevents other people from making money off of your work. For the rest of your life, plus 75 years.

Something is broken here, alright. While we're discussing intellectual property, what about one's DNA? Is it not a performance of biology? How about your fingerprint? Fingerprints are semi-unique, so it's also a performance mark. We've seen celebrities sue for the use of their likeness, so that's recognized to some degree as well. At what point will data subjects get rights, so that when another Equifax happens, they can be bankrupted and prevented from harming the public again?

Most are rhetorical, of course, but I really think generative language models are disrupting a lot of things we used to take for granted, and our models of creatorship are not refined enough to account for digital or statistical copying. Intellectual property as a concept is not compatible with a digital future.

Post reply on HN