Live data from Hacker News

New York Times considers legal action against OpenAI as copyright tensions swirl

npr.org

171–180 of 383 posts

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#171
post #142

Earlier quoted context omitted.

> we end up with an inferior product cannibalizing a superior one and driving it out of business. In case that print is meant by inferior product: The same argument could've been brought up for Napster, where traditional distribution via CD printing through music labels are the inferior product driving the superior one out of business. Or rather it's big labels suing Napster out of business. I also hold a dislike for…

> A ruling in favor of copyright would force OpenAI to shut down - but given their impressive demo of the tech Who cares? If they want data, they can pay for it.

Or at least ask before scrapping/reading it.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#172

Don't humans operate similarly? We gain knowledge through experiences. These AI models effectively condense a vast amount of experience data into weights. Considering the global race in AI advancements, I'm skeptical about the success of these copyright claims. I do find it hypocritical that OpenAI says that other LLMs can't be trained on data generated by their LLMs.

Yes, humans operate similarly. We gain knowledge through learning and assimilation of experiences. But there is a cost associated with each new input, in one form or another. You can read and reproduce the NYT articles or their content and build based on what you have learned. But aside from piracy, you have to somehow pay a cost associated with access (ads, subscriptions etc). The LLM does not do so in an equal way.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#173

Don't humans operate similarly? We gain knowledge through experiences. These AI models effectively condense a vast amount of experience data into weights. Considering the global race in AI advancements, I'm skeptical about the success of these copyright claims. I do find it hypocritical that OpenAI says that other LLMs can't be trained on data generated by their LLMs.

> Don't humans operate similarly? I'm going to bypass this question a bit and say, who cares? Why do we need to treat these things the same way we treat humans? Why can we not say that it's okay if a human does it, and not okay if it's a computer? There's nothing that requires us to establish 'fair' as treating them the same as people.

> Why do we need to treat these things the same way we treat humans?

Because it would be absurd for it to be legal to do X, but illegal to do so with an efficient tool. Especially when the activity X in question is "learning".

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#175
post #154

Earlier quoted context omitted.

Yup, and that reproducing it is a copyright violation. Still a copyright abolitionist though. Maybe now more people will join the fight?

Well if use the tool to reproduce copyrighted content you are violating copyright. But that’s not the primary usecase and nobody in their right mind is arguing that. The weights are not a reproduction of the content. They are capable of it but so is a photocopier a lot more and we didn’t ban those either despite them technically being a lot more useful for violation. Nah, this is expansionist doctrine and agenda for…

Photocopiers are for personal use, training an AI is not. If you photocopy 10,000 copies of copyrighted text and starting distributing it you will get sued.

It would be different if I trained my own AI, for my personal use.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#176

I think that the proper outcome for all of this would be acknowledgement that the current copyright laws very poorly regulate this aspect, that the key parts of any such legal action are at the not-really-described edges of law because these edges weren't relevant until now; and so instead of waiting for courts ruling on how law-as-written-now applies and accepting these rulings, we will likely get some new legislati…

100% this. But I doubt it will happen in the US, unfortunately.

The tech industry has sufficient money and influence for lobbying to push this one through. The media industry did the DMCA adjustments to copyright reasonably fast, and tech industry is even more powerful and wealthy.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#177

IANAL, but copyright protections are pretty much tied to content and format and not to the idea itself, with the intent of preventing (or putting a price on) the copying of original works. The Times will have a very hard time proving that their content is being re-marketed by OpenAI. Having a competing product based on your ideas. Compare: "Steve Jobs [was] a tyrant": https://www.nytimes.com/2011/10/07/technology/ste…

I think this is going to be a test of the fair use doctrine.

https://www.copyright.gov/fair-use/

Now, there's this idea that "news" is just factual and therefore falls under "fair use". However, that's only part of what section 107 says.

Fair use very much is still conditional, as there are 4 factors to be considered: (a) Purpose and character of the use, including whether the use is of a commercial nature or is for nonprofit educational purposes (b) Nature of the copyrighted work (c) Amount and substantiality of the portion used in relation to the copyrighted work as a whole and (d) Effect of the use upon the potential market for or value of the copyrighted work

The big issue isn't companies training LLM's using unlicensed materials (e.g. copyright protected works); it's publishing the output to the wider world. That's where a liability is created.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#178

Earlier quoted context omitted.

> Don't humans operate similarly? I'm going to bypass this question a bit and say, who cares? Why do we need to treat these things the same way we treat humans? Why can we not say that it's okay if a human does it, and not okay if it's a computer? There's nothing that requires us to establish 'fair' as treating them the same as people.

> Why do we need to treat these things the same way we treat humans? Because it would be absurd for it to be legal to do X, but illegal to do so with an efficient tool. Especially when the activity X in question is "learning".

Why would it be absurd?

Did we somehow stealthily develop a neural interface that lets us feed the 'learning' that 'AI' is doing into a human brain? Have we actually figured out how to do that?

No, we haven't. So humans are still learning the same way, but with a new tool to condense and summarize some information. Kinda like a textbook in school. But we don't treat those as human beings with human rights do we?

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#179

Don't humans operate similarly? We gain knowledge through experiences. These AI models effectively condense a vast amount of experience data into weights. Considering the global race in AI advancements, I'm skeptical about the success of these copyright claims. I do find it hypocritical that OpenAI says that other LLMs can't be trained on data generated by their LLMs.

Computers aren't humans and LLMs aren't human brains. We have no way to reconstruct memories from a preserved brain (yet). The exact ways in which humans form memories and store information isn't even known yet; we're still drilling into the specifics from higher-level concepts. Modeling the human brain like nodes with weights ignores a lot of biological processes. Blood/oxygen flow, hormones, neurotransmitter decay,…

So if we add a bunch of complex processes to an LLM, in order to produce a better analog of a human in terms of degree of complexity if not actual function, does that have some bearing on this copyright question?

It doesn’t seem clear to me that it does.

Is the argument that sufficient complexity in how an “intelligence” processes this copyrighted data leads to the output being transformative vs not transformative in a more simple mind/model?

What if the output is exactly the same, or comparable enough, regardless of the degree of complexity of the mind/model?

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#180
Hopefully soon enough (within a decade?) we’ll all be able to run large language models on cheap consumer devices, and model weights containing everything including NYT will be floating around in the form of warez readily consumed by anyone with a modicum of savvy, whether NYT likes them or not. They can’t stop progress.
Post reply on HN