Live data from Hacker News

OpenAI now tries to hide that ChatGPT was trained on copyrighted books

businessinsider.com

41–50 of 94 posts

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#41
post #3

I think the conclusion is slightly wrong: they're not trying to hide the training process. They'll probably wind up vigorously defending that in court however they can, they're in big trouble if they can't train like that. They're trying to avoid reproducing copyrighted text, which is a totally separate (and arguably more clear-cut) legal question. Input vs output.

> Input vs output. This is what I've been wondering. Does Fair Use apply here at all? Sure, the models were trained on copyrighted material. But wouldn't the generative part of the AI count as transformative?

I was trained on copyrighted books. I only get in trouble when I spout out paragraphs from them from memory and pass them off as my own.

I know not to do that, though.

Seems only fair that the same should apply to GPT.

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#42
post #13

People keep comparing it to a human ingesting content throughout their life and then being influenced in their own works. I'm sorry but that is not the same thing - at all. The concept of training a model with the explicit intent of selling the output of that model is inherently different. Not saying that it should be illegal. But it is clearly in violation of the spirit of existing copyright law, in my opinion. They…

Please explain how this is any different than the Google Books case. Please do so without any feelings involved to the best of your ability.

Even if OpenAI maliciously ingested copywritten work, it is just a bunch of numbers and if you go in and ask for it spit the book back out, it won't. That simply isn't how this works.

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#43
The article's focus on copyright is inflammatory, but the paper does present some big questions.

Original paper: https://huggingface.co/papers/2308.05374

Abstract:

  major categories of LLM trustworthiness: 
   reliability, safety, fairness, resistance to misuse, 
   explainability and reasoning, adherence to social norms, and robustness. 

  Each major category is further divided into [...] a total of 29 sub-categories. 

   The measurement results indicate that, in general, 
   more aligned models tend to perform better in terms of overall trustworthiness
The paper shows that trustworthiness is now a design goal. It would seem good to avoid e.g., hallucinations, but is it really socially good to induce reliance?

Compare AI companies targeting "trustworthiness" today with social media companies targeting "engagement" in decades past: they might increase uptake, but build significant externalities into their business model, and then lose most of the social benefits of adoption.

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#44
post #32
post #31

Earlier quoted context omitted.

Are you trolling? Look at the ISBN/info page of literally any book. It will say something like 'all rights reserved' and that you can't reproduce any part without permission, except for limited cases. Please go get a book off of your shelf and look, I'm begging you. Edit - Example from Infinite Jest: https://burnsiderarebooks.cdn.bibliopolis.com/pictures/14094...

Using the knowledge from the book !== reproducing any part of the book. If I read a math book that shows me how to do an integral, then use that knowledge to do integrals, I'm not infringing the copyright of the book, ffs. If you read in a book that the main export of Germany is Bavarian creme doughnuts, and you use that knowledge in a job interview (i.e. making money) to land a job as a Bavarian creme doughnut impor…

I am not saying that using knowledge is a copyright violation. Sorry but I think we're talking past each other.

Inputting entire books verbatim into a model is not the same as a human learning a fact and using it later.

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#46
post #13

People keep comparing it to a human ingesting content throughout their life and then being influenced in their own works. I'm sorry but that is not the same thing - at all. The concept of training a model with the explicit intent of selling the output of that model is inherently different. Not saying that it should be illegal. But it is clearly in violation of the spirit of existing copyright law, in my opinion. They…

> The concept of training a model with the explicit intent of selling the output of that model is inherently different

Uh-oh, better tell the colleges and universities to stop promoting degree programs off the back of the potential increase in lifetime earnings. Wouldn't want anyone to get the idea that it's okay to absorb a bunch of copyrighted material during college to update their mental model, then make money by selling the output of that model for the rest of their lives.

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#48
post #26

Earlier quoted context omitted.

It is not illegal to make money with copyrighted content. Nor should it be

But did you not purchase that content first?

but buying a book is not buying the copyright of it.

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#49
post #13

People keep comparing it to a human ingesting content throughout their life and then being influenced in their own works. I'm sorry but that is not the same thing - at all. The concept of training a model with the explicit intent of selling the output of that model is inherently different. Not saying that it should be illegal. But it is clearly in violation of the spirit of existing copyright law, in my opinion. They…

Yes, it is different that human has magic and machine don't. I learn stuff from books and sell my skills. I definitely copied someone's knowledge about calculus, engineering, and I am intented to earn money from them. Machine cannot do this because only me has magic.

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#50
post #31

Earlier quoted context omitted.

Are you trolling? Look at the ISBN/info page of literally any book. It will say something like 'all rights reserved' and that you can't reproduce any part without permission, except for limited cases. Please go get a book off of your shelf and look, I'm begging you. Edit - Example from Infinite Jest: https://burnsiderarebooks.cdn.bibliopolis.com/pictures/14094...

And where does it say "you can't learn anything from this book at all". We're not talking about copying verbatim, that's the whole point.

But we are talking about copying verbatim. The entire text didn't just land in the training corpus by magic.

A human copied the entire text of many books - verbatim - into the training corpus.

Under copyright law, the human was free to read the whole thing. But not copy the whole thing into an AI model with the intent to profit.

Post reply on HN