Live data from Hacker News

OpenAI now tries to hide that ChatGPT was trained on copyrighted books

businessinsider.com

81–90 of 94 posts

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#81

it would seem to me that from a technical perspective the weights of an AI model trained on copyrighted material would be a reproduction of the copyrighted work. just because you combine the information from millions (or more) copyrighted works together doesn't mean that you aren't reproducing them. the process of training requires reproduction and distribution of the works internally as part of the data processing p…

Distributing data to servers only accessible by a machine isn't what folks have in mind when they talk about the distribution of copyrighted content.

> isn't what folks have in mind

you would have to ask a judge about that because its a novel legal question. obviously nobody anticipated this technology at the time it was written so ultimately it will have to be a court that decides how to apply existing laws.

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#82
post #36

it would seem to me that from a technical perspective the weights of an AI model trained on copyrighted material would be a reproduction of the copyrighted work. just because you combine the information from millions (or more) copyrighted works together doesn't mean that you aren't reproducing them. the process of training requires reproduction and distribution of the works internally as part of the data processing p…

If I publish an article on a subject I extensively read about on books and add no new information, I’m just reproducing them. Should that be considered a violation of copyright?

false analogy. if openai is downloading and copying and distributing copyrighted material internally and then compressing that information into a model that can reproduce it later and selling access to that model that is a very different thing.

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#83

Earlier quoted context omitted.

You're needlessly complicating this and also changing what the parent comment I was responding do takes issue with. There are two things to be evaluated here: 1) Is ingestion of copyrighted material as part of the model training process an infringement itself? I say no - it is equivalent to a human going to a library and reading all the books there. The knowledge gained by the human is equivalent to the weights that…

> Is ingestion of copyrighted material as part of the model training process an infringement itself? I say no - it is equivalent to a human going to a library and reading all the books there. It might be Fair Use for other reasons, but any argument that uses an “it’s like a human doing X” analogy is, legally, misguided. The courts simply do not see machines as being legally analogous to human brains. Humans can “copy…

[deleted]

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#84
post #79

Earlier quoted context omitted.

I sure didn't buy a license for every book and screenplay I've ever read. And I borrow a lot of books from the library. I'm not understanding why you think ML models need to license all of their input content when humans clearly do not.

If you bought the books and screenplays, you acquired them in a way that doesn't infringe on copyright. If you borrowed them from the library, you acquired them in a way that doesn't infringe on copyright. You don't need a license because you aren't republishing. ML training doesn't work without having a copy of the data (books and screenplays). That data can either be copied in a non-infringing way (buy the books; a…

[deleted]

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#85
post #30

Earlier quoted context omitted.

I'm not understanding how it's different. For example, if I specifically set out to make money by creating and selling a parody of Harry Potter by reading all the Harry Potter books a bunch of times, does that make my parody a violation of the spirit of existing copyright law? Edit to add another example because someone is going to say that parody is its own thing and exempted. If I want to make money by writing a fi…

Did you, through copyright infringement, acquire your copies of Harry Potter? Did you use Harry Potter source material and a machine to transform the input and produce your parody material? I think it's different if OpenAI acquires licenses for all of the material it uses for training.

Doesn't matter. The only thing copyright regulates is the duplication and spreading of copyrighted material. If you acquire "unlicensed" copies and learn from them, you have not violated copyright, you are not violating the author's copyright.

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#86

Earlier quoted context omitted.

> The concept of training a model with the explicit intent of selling the output of that model is inherently different Uh-oh, better tell the colleges and universities to stop promoting degree programs off the back of the potential increase in lifetime earnings. Wouldn't want anyone to get the idea that it's okay to absorb a bunch of copyrighted material during college to update their mental model, then make money by…

You do understand that those are humans we are talking about. Humans with lives and livelihoods and wants and needs. Wishing to learn for the sake of learning and growing as well as sustaining themselves? And that openai is a company trying to make money? You do see the enormous difference there I hope? Let alone the huge shift of power from one to many (one university to many students), to many to one (many resource…

I don’t get the ‘OpenAI is a company not a person’ argument. Would those arguments go away if the GPT model had been published by a single individual, Satoshi style? The question of whether it’s okay for OpenAI to train an LLM on copyrighted material and sell the results would also go to whether you or I can train an LLM (or finetune one, or do any other kind of ML training), on copyrighted material and make commercial use of the result.

It seems to me OpenAI have done the equivalent of distilling a bunch of knowledge down into a book. And they’ve given that book a really good index.

They’re essentially an encyclopedia vendor. Just an exceedingly sophisticated one.

Back in the olden times publishers used to pay people to write encyclopedia entries based on summarizing stuff they had read in other books. Nowadays we still do the same thing only with volunteer time (Wikipedia requires everything it contains to be externally sourced, after all. You can’t write anything into a Wikipedia article without reading it somewhere copyrighted first).

OpenAI’s model isn’t so different. Source material, indexed and summarized to make it easier to search and use.

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#87
post #36

Earlier quoted context omitted.

If I publish an article on a subject I extensively read about on books and add no new information, I’m just reproducing them. Should that be considered a violation of copyright?

false analogy. if openai is downloading and copying and distributing copyrighted material internally and then compressing that information into a model that can reproduce it later and selling access to that model that is a very different thing.

What if I just make a robot to go to public library and book store and then OCR everything, would it make a difference?

Again there is no "compression of information" in deep learning.

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#88
post #36

Earlier quoted context omitted.

If I publish an article on a subject I extensively read about on books and add no new information, I’m just reproducing them. Should that be considered a violation of copyright?

The analogy would hold if a single human could read all books and then be copied an infinite number of times with little effort and respond to an infinite number of prompts simultaneously. This is fundamentally different than the example of a single person reading books and being inspired by them to produce something. The scale is important and I see this lost in a lot of discussions.

Simply no. Logic must be consistent from 1 to infinity. If someone read and speak too fast is illegal, then let's do it. Put everyone has IQ > 130 into jail. Oh Rainman cannot forget things, obviously that is illegal, shot it down.

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#89
post #30

Earlier quoted context omitted.

Did you, through copyright infringement, acquire your copies of Harry Potter? Did you use Harry Potter source material and a machine to transform the input and produce your parody material? I think it's different if OpenAI acquires licenses for all of the material it uses for training.

Doesn't matter. The only thing copyright regulates is the duplication and spreading of copyrighted material. If you acquire "unlicensed" copies and learn from them, you have not violated copyright, you are not violating the author's copyright.

You seem to be arguing that downloading copyrighted material without permission is not copyright infringement. "Acquire 'unlicensed' copies" is exactly downloading without permission. You are making a copy, thus infringing on the author's copyright.
Post reply on HN