Live data from Hacker News

OpenAI now tries to hide that ChatGPT was trained on copyrighted books

businessinsider.com

51–60 of 94 posts

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#51
post #35

Earlier quoted context omitted.

I'm not understanding how it's different. For example, if I specifically set out to make money by creating and selling a parody of Harry Potter by reading all the Harry Potter books a bunch of times, does that make my parody a violation of the spirit of existing copyright law? Edit to add another example because someone is going to say that parody is its own thing and exempted. If I want to make money by writing a fi…

Surely you must understand the difference between copying something and producing something entirely new. If you literally only consumed Tarantino media and had no other influences, and then one day made a movie so great everyone was calling you the next Tarantino - that would be fine! As long as you didn't copy anything that he actually made himself. Being influenced by is not the same as copying.

Of course Tarantino copied a ton for his films…

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#52

it would seem to me that from a technical perspective the weights of an AI model trained on copyrighted material would be a reproduction of the copyrighted work. just because you combine the information from millions (or more) copyrighted works together doesn't mean that you aren't reproducing them. the process of training requires reproduction and distribution of the works internally as part of the data processing p…

I trust this just because they trust they can escape from this.

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#53
post #13

People keep comparing it to a human ingesting content throughout their life and then being influenced in their own works. I'm sorry but that is not the same thing - at all. The concept of training a model with the explicit intent of selling the output of that model is inherently different. Not saying that it should be illegal. But it is clearly in violation of the spirit of existing copyright law, in my opinion. They…

Yes, it is different that human has magic and machine don't. I learn stuff from books and sell my skills. I definitely copied someone's knowledge about calculus, engineering, and I am intented to earn money from them. Machine cannot do this because only me has magic.

why computer in openAI have magic can let author lose their copyright ? you know, because they have a magical algorithm?

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#54
post #35

Earlier quoted context omitted.

Surely you must understand the difference between copying something and producing something entirely new. If you literally only consumed Tarantino media and had no other influences, and then one day made a movie so great everyone was calling you the next Tarantino - that would be fine! As long as you didn't copy anything that he actually made himself. Being influenced by is not the same as copying.

But the whole point is that copying is not what's happening .

Copyrighted content was not _copied_ into the training corpus?

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#55
post #44
post #32

Earlier quoted context omitted.

Using the knowledge from the book !== reproducing any part of the book. If I read a math book that shows me how to do an integral, then use that knowledge to do integrals, I'm not infringing the copyright of the book, ffs. If you read in a book that the main export of Germany is Bavarian creme doughnuts, and you use that knowledge in a job interview (i.e. making money) to land a job as a Bavarian creme doughnut impor…

I am not saying that using knowledge is a copyright violation. Sorry but I think we're talking past each other. Inputting entire books verbatim into a model is not the same as a human learning a fact and using it later.

The book does not exist verbatim within the model. You cannot open the model and find the full text of some copyrighted work anywhere.

The model just trains a set of weights for a neural network, then probabilistically generates text in response to prompts.

And yeah, humans "input entire books verbatim" (aka, reading) and then regurgitate that knowledge later when they determine it is likely to be an appropriate response to a question or comment from someone else. It's really not that different. It doesn't even have to be a fact. How many responses on the Internet are just quotes from movies or TV shows?

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#60
post #13

People keep comparing it to a human ingesting content throughout their life and then being influenced in their own works. I'm sorry but that is not the same thing - at all. The concept of training a model with the explicit intent of selling the output of that model is inherently different. Not saying that it should be illegal. But it is clearly in violation of the spirit of existing copyright law, in my opinion. They…

Please explain how this is any different than the Google Books case. Please do so without any feelings involved to the best of your ability. Even if OpenAI maliciously ingested copywritten work, it is just a bunch of numbers and if you go in and ask for it spit the book back out, it won't. That simply isn't how this works.

I don't see much of a difference, seems like google profited off of copyrighted material that they didn't own.

I dont understand the purpose of your feelings comment. I have no affiliations or preferences for any dogs in this fight.

I don't necessarily think they are identical scenarios but if I were OpenAI's lawyer that's probably what I'd try and point at.

I didnt say it would spit the book back out, at any point. I've said a human chose to copy the entirety of copyrighted texts, verbatim, into the training corpus for a language model that they intended to sell.

Post reply on HN