Live data from Hacker News

New York Times considers legal action against OpenAI as copyright tensions swirl

npr.org

101–110 of 383 posts

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#101

Earlier quoted context omitted.

But the copy used during training is itself transient, so all this boils down really to the question of whether training a machine is fair use. Which can't be answered here exactly because the concept of fair use is deliberately vague, so this will boil down to a lawsuit and probably go to the Supremes. The USA will work something out that's reasonable as they always do and, lacking AI companies and often the concept…

It's not even a question of "fair use" because OpenAI isn't providing anybody a copy. Probability models really just aren't copyrightable to begin with.

Fair use also covers generating derived works, and arguably the AI is a derived work.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#102

Earlier quoted context omitted.

Really? Seems like it's not so different conceptually to search engines. They also make money by indexing "training data" of a sort. But the open web trundles on regardless, because the search engines found ways to cut the content producers in on it. I see no reason why that can't be the case here too, with AI companies training their models to act more like search engines when data comes from certain sources - i.e.…

>Really? Seems like it's not so different conceptually to search engines. They also make money by indexing "training data" of a sort. OK, so if a writer X has a blog to put up samples of their work to drive people to buy books and to get writing assignments and someone uses ChatGPT to write something in the style of X - this naively seems like a hit on that author's ability to sell their skills. And I don't think it…

Seems like a rerun of the argument over snippets, or the "answer onebox" as Google used to call it where the info you need is directly inlined in the SERP rather than being behind a link.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#103

Earlier quoted context omitted.

But the copy used during training is itself transient, so all this boils down really to the question of whether training a machine is fair use. Which can't be answered here exactly because the concept of fair use is deliberately vague, so this will boil down to a lawsuit and probably go to the Supremes. The USA will work something out that's reasonable as they always do and, lacking AI companies and often the concept…

It's not even a question of "fair use" because OpenAI isn't providing anybody a copy. Probability models really just aren't copyrightable to begin with.

> It’s not even a question of “fair use” because OpenAI isn’t providing anybody a copy

“Fair Use” can apply to essentially any of the exclusive rights under copyright, including making (with or without distributing) a copy or derivative work.

> Probability models really just aren’t copyrightable to begin with.

“Probability models” that are built from copyrightable works either:

(1) have a sufficient human creative input to be copyrightable on its own (in which caee it still may infringe copyrights applicable to its source material as a derivative work), or

(2) do not have a sufficient human creative input to be copyrightable, and thus are a form of mechanical copy of their data from which they are developed (which, to the extent it either is, in aggregate, a copyrighted work, and/or contains copies of other copyrighted works, is protected by one or more copyrights, which do not cease to apply to the mechanical copy, and which an unlicensed mechanical copy would violate unless it fell into an exception like Fair Use.)

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#105

Earlier quoted context omitted.

It's not even a question of "fair use" because OpenAI isn't providing anybody a copy. Probability models really just aren't copyrightable to begin with.

> It’s not even a question of “fair use” because OpenAI isn’t providing anybody a copy “Fair Use” can apply to essentially any of the exclusive rights under copyright, including making (with or without distributing) a copy or derivative work. > Probability models really just aren’t copyrightable to begin with. “Probability models” that are built from copyrightable works either: (1) have a sufficient human creative in…

But OpenAI is neither copying nor deriving a work. Style (which can be described as a probability model) is not copyrightable.

You're halfway there with #2. The output is not copyrightable, but unless you can actually point to a sequence of words from the original it can't be infringing.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#106

Earlier quoted context omitted.

It's not even a question of "fair use" because OpenAI isn't providing anybody a copy. Probability models really just aren't copyrightable to begin with.

Fair use also covers generating derived works, and arguably the AI is a derived work.

That would be hard to argue. You can't point to a sequence of words that's in the original that was copied into the OpenAI output.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#107
post #70

There is a very real risk that we end up with an inferior product cannibalizing a superior one and driving it out of business. Moreover, AI would seem to be even more susceptible to capture and manipulation than conventional media. When it's a question of guiding thought I prefer the humanities to tech. (Same with art.)

> we end up with an inferior product cannibalizing a superior one and driving it out of business. In case that print is meant by inferior product: The same argument could've been brought up for Napster, where traditional distribution via CD printing through music labels are the inferior product driving the superior one out of business. Or rather it's big labels suing Napster out of business. I also hold a dislike for…

For me the question is highly debatable. As far as I'm aware the training of AI works by crawling various content from the Internet, so they're using a product, which are NY articles to train an AI, meaning they're using their content to help creating a product.

But in this sense, shouldn't they be suing Google as well? Since Google as a search engine, also crawls the web and shows their articles in their search results, usually it may even use them for those quick answers features.

My 5 cents on this are that NY Times noticed OpenAI has deep pockets, they may have ground to sue and decides to try their look in order to get some quick easy money. Now, I don't know if what OpenAI is doing with ChatGPT does not fall under fair use.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#108
Don't humans operate similarly? We gain knowledge through experiences. These AI models effectively condense a vast amount of experience data into weights. Considering the global race in AI advancements, I'm skeptical about the success of these copyright claims. I do find it hypocritical that OpenAI says that other LLMs can't be trained on data generated by their LLMs.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#109

Copyrights and patents are holding back humanity.

Somebody posted this link in a comment on a thread the other day: https://news.artnet.com/market/koch-brother-loses-it-on-air-...

And it occurred to me that this is precisely the thing that's holding back humanity: "Koch estimates that he has spent $25 million on legal fees—far more than the $5 million he originally spent on the fake wine itself."

We have a legal system that is completely inaccessible to the average man. That it can generate $25M in civil legal fees is beyond absurd. Patents and copyrights derive much of their force from the fact that they are enforced by a legal process where to play is to lose. There's no winning. It's no longer about justice, and it has largely become a form of financial bullying where entrenched interests beat up on smaller ones.

Fix the legal system -- make it accessible -- and you've fixed patents and copyrights. To address this thread's point: I believe that AI, and perhaps _only_ AI, might be able to help with this.

Post reply on HN