Live data from Hacker News

New York Times considers legal action against OpenAI as copyright tensions swirl

npr.org

201–210 of 383 posts

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#201

Don't humans operate similarly? We gain knowledge through experiences. These AI models effectively condense a vast amount of experience data into weights. Considering the global race in AI advancements, I'm skeptical about the success of these copyright claims. I do find it hypocritical that OpenAI says that other LLMs can't be trained on data generated by their LLMs.

Humans can be sued for plagiarism, so AI should be as well.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#202
post #184

Earlier quoted context omitted.

So if we add a bunch of complex processes to an LLM, in order to produce a better analog of a human in terms of degree of complexity if not actual function, does that have some bearing on this copyright question? It doesn’t seem clear to me that it does. Is the argument that sufficient complexity in how an “intelligence” processes this copyrighted data leads to the output being transformative vs not transformative in…

IMHO no because LLM's are not legal persons and most likely will never be, so they can't acquire or operate under laws, privileges and agreements simply because they exists.

Yeah I agree that legal personhood for LLMs at this point is far-fetched.

This would be a separate argument, though, from the notion I responded to above that the difference in the processes of a human mind and LLM are the reason why "learning" from copyrighted material is a violation of copyright in one case and not the other.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#203

Earlier quoted context omitted.

In order to render that page, it probably was copied dozens of times all over my RAM. Do I owe NYT money now?

Well, are you making money off of those copies?

Not personally, but if I was training to be a journalist then yes.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#204

IANAL, but copyright protections are pretty much tied to content and format and not to the idea itself, with the intent of preventing (or putting a price on) the copying of original works. The Times will have a very hard time proving that their content is being re-marketed by OpenAI. Having a competing product based on your ideas. Compare: "Steve Jobs [was] a tyrant": https://www.nytimes.com/2011/10/07/technology/ste…

I tend to agree with you, but, one could argue “statistical collection of words” is a form of compression? For example, you can’t write a kids version of a novel and sell that without dealing with copyright.

The part openai will have to argue is that it's not mererly compression but an irreversible transformation.

Which is hard, best hope they have is trying to put the burden of proof on the nytimes to show you can make the model regurgitate their articles (with some nudging).

If they manage that then nytimes is going to have a lot of trouble showing the model actually breaches their copyright, because just the information contained in their articles is not enough to constitute a copyrightable work.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#205
post #141

Don't humans operate similarly? We gain knowledge through experiences. These AI models effectively condense a vast amount of experience data into weights. Considering the global race in AI advancements, I'm skeptical about the success of these copyright claims. I do find it hypocritical that OpenAI says that other LLMs can't be trained on data generated by their LLMs.

It's quite different, is it not? I don't get the analogy. These models are scanning, storing, and ingesting more material than any one human could. Not only is the method completely different, the end goal and applications are as well. The analogy basically isn't one, at all. I'm pretty upset at companies using our personal data to make gobs of money off of. I'm also upset that they're now using our knowledge work to…

>It's quite different, is it not? I don't get the analogy. These models are scanning, storing, and ingesting more material than any one human could. Not only is the method completely different, the end goal and applications are as well. The analogy basically isn't one, at all.

So you'd be fine with it if the model only ingested a humanly plausible amount of data? I suspect that would only make their legal issues worse, since the LLM would be much more likely to repeat tokens from the training set verbatim.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#206

Earlier quoted context omitted.

In order to render that page, it probably was copied dozens of times all over my RAM. Do I owe NYT money now?

Well, are you making money off of those copies?

Does reading news articles not benefit you? If not why continue reading?

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#207
post #196

Earlier quoted context omitted.

Photocopiers are for personal use, training an AI is not. If you photocopy 10,000 copies of copyrighted text and starting distributing it you will get sued. It would be different if I trained my own AI, for my personal use.

businesses use photocopiers, there’s Xerox shops, etc. Again, LLMs don’t copy so it’s not a good metaphor.

I know. It's well established.

You brought up photocopiers and you never said it wasn't a good metaphor so I don't know why you pre-pended "Again".

If it's not a good metaphor, that invalidates your point that:

"we didn’t ban [photocopiers] either despite them technically being a lot more useful for violation"

So now you're arguing against yourself.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#208
post #54

Earlier quoted context omitted.

Are you using this knowledge to produce a product that materially undercuts NYT revenue?

That doesn’t matter though. If I licensed material why should it count if I compete or not. If I watch Lebron James play and use it to develop an athletic training program, does it matter if I play baseball? Or WNBA? Or does it only matter if I play against him in the championship?

Licenses are bound to the purposes specified by the license. So go read the fine print.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#209
post #90

> if a federal judge finds that OpenAI illegally copied The Times' articles to train its AI model, the court could order the company to destroy ChatGPT's dataset, forcing the company to recreate it using only work that it is authorized to use. I'd like to see it happening but it sounds unrealistic.

If I read 1000s of of NYT articles to improve my writing skills, add then write an article of my own, is that a copyright violation?

The issue is that you know well enough where the dividing line is between a fully original article you write and one that plagiarizes. LLMs on the other hand don’t have that awareness, and can plagiarize by accident.

Re: New York Times considers legal action against OpenAI as copyright tensions swirl

#210

Earlier quoted context omitted.

Or on the converse: if those industries are unviable without copyright protection, they could go away entirely. This is a plausible path to "drop copyright entirely", just like encryption was dropped as an export-controlled technology in the late 90s. (remember the 40-bit "international" SSL?) OpenAI etc. have huge amounts of money behind them, they very well have a fighting chance in court to defend their usage of s…

> if those industries are unviable without copyright protection, they could go away entirely. These creative industries include all of software development, music, TV, movies, books, media, art, etc. You do technically solve the problem of copyright by shutting all those down, but I'm not sure it's a solution anybody will vote for. If you can come up with a serious alternative though, which can sustain those creative…

I don't think that no one would create new software/music/books/movies/art/etc. without copyright. Humans have done so for millenia before, they still did so in absence of copyright protections. I don't see how this is not a serious alternative - the only major losers would be the middlemen, not the artists themselves.
Post reply on HN