Live data from Hacker News

OpenAI now tries to hide that ChatGPT was trained on copyrighted books

businessinsider.com

31–40 of 94 posts

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#31
post #27
post #23

Earlier quoted context omitted.

Those people are entering into and abiding by a license agreement when they purchase the book.

No they aren't. Where's the Terms of Service for a book? What about secondhand books or books you get from the library or books you borrow from a friend? I never had a book tell me to an accept a license agreement before I could read it.

Are you trolling? Look at the ISBN/info page of literally any book.

It will say something like 'all rights reserved' and that you can't reproduce any part without permission, except for limited cases.

Please go get a book off of your shelf and look, I'm begging you.

Edit - Example from Infinite Jest: https://burnsiderarebooks.cdn.bibliopolis.com/pictures/14094...

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#32
post #31
post #27

Earlier quoted context omitted.

No they aren't. Where's the Terms of Service for a book? What about secondhand books or books you get from the library or books you borrow from a friend? I never had a book tell me to an accept a license agreement before I could read it.

Are you trolling? Look at the ISBN/info page of literally any book. It will say something like 'all rights reserved' and that you can't reproduce any part without permission, except for limited cases. Please go get a book off of your shelf and look, I'm begging you. Edit - Example from Infinite Jest: https://burnsiderarebooks.cdn.bibliopolis.com/pictures/14094...

Using the knowledge from the book !== reproducing any part of the book.

If I read a math book that shows me how to do an integral, then use that knowledge to do integrals, I'm not infringing the copyright of the book, ffs.

If you read in a book that the main export of Germany is Bavarian creme doughnuts, and you use that knowledge in a job interview (i.e. making money) to land a job as a Bavarian creme doughnut importer, that is not copyright infringement. People learn things and put that knowledge to use, often for profit.

What kind of insane interpretation of copyright law are you working with?

A copyright notice is not an agreement. It's not a contract, it does not offer any consideration to the other party. There is no meeting of the minds. That is not at all how any of that works.

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#33
it would seem to me that from a technical perspective the weights of an AI model trained on copyrighted material would be a reproduction of the copyrighted work. just because you combine the information from millions (or more) copyrighted works together doesn't mean that you aren't reproducing them.

the process of training requires reproduction and distribution of the works internally as part of the data processing pipeline so why wouldn't you need a license for that?

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#34
post #13

People keep comparing it to a human ingesting content throughout their life and then being influenced in their own works. I'm sorry but that is not the same thing - at all. The concept of training a model with the explicit intent of selling the output of that model is inherently different. Not saying that it should be illegal. But it is clearly in violation of the spirit of existing copyright law, in my opinion. They…

It is not illegal to make money with copyrighted content. Nor should it be

The principal criminal statute protecting copyrighted works is 17 U.S.C. § 506(a), which provides that "[a]ny person who infringes a copyright willfully and for purposes of commercial advantage or private financial gain" shall be punished as provided in 18 U.S.C. § 2319. Section 2319 provides, in pertinent part, that a 5-year felony shall apply if the offense "consists of the reproduction or distribution, during any 180-day period, of at least 10 copies or phonorecords, of 1 or more copyrighted works, with a retail value of more than $2,500." 18 U.S.C. § 2319(b)(1).

https://www.justice.gov/archives/jm/criminal-resource-manual...

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#35
post #13

People keep comparing it to a human ingesting content throughout their life and then being influenced in their own works. I'm sorry but that is not the same thing - at all. The concept of training a model with the explicit intent of selling the output of that model is inherently different. Not saying that it should be illegal. But it is clearly in violation of the spirit of existing copyright law, in my opinion. They…

I'm not understanding how it's different. For example, if I specifically set out to make money by creating and selling a parody of Harry Potter by reading all the Harry Potter books a bunch of times, does that make my parody a violation of the spirit of existing copyright law? Edit to add another example because someone is going to say that parody is its own thing and exempted. If I want to make money by writing a fi…

Surely you must understand the difference between copying something and producing something entirely new.

If you literally only consumed Tarantino media and had no other influences, and then one day made a movie so great everyone was calling you the next Tarantino - that would be fine!

As long as you didn't copy anything that he actually made himself.

Being influenced by is not the same as copying.

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#36

it would seem to me that from a technical perspective the weights of an AI model trained on copyrighted material would be a reproduction of the copyrighted work. just because you combine the information from millions (or more) copyrighted works together doesn't mean that you aren't reproducing them. the process of training requires reproduction and distribution of the works internally as part of the data processing p…

If I publish an article on a subject I extensively read about on books and add no new information, I’m just reproducing them. Should that be considered a violation of copyright?

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#37
post #35

Earlier quoted context omitted.

I'm not understanding how it's different. For example, if I specifically set out to make money by creating and selling a parody of Harry Potter by reading all the Harry Potter books a bunch of times, does that make my parody a violation of the spirit of existing copyright law? Edit to add another example because someone is going to say that parody is its own thing and exempted. If I want to make money by writing a fi…

Surely you must understand the difference between copying something and producing something entirely new. If you literally only consumed Tarantino media and had no other influences, and then one day made a movie so great everyone was calling you the next Tarantino - that would be fine! As long as you didn't copy anything that he actually made himself. Being influenced by is not the same as copying.

But the whole point is that copying is not what's happening.

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#38

it would seem to me that from a technical perspective the weights of an AI model trained on copyrighted material would be a reproduction of the copyrighted work. just because you combine the information from millions (or more) copyrighted works together doesn't mean that you aren't reproducing them. the process of training requires reproduction and distribution of the works internally as part of the data processing p…

Distributing data to servers only accessible by a machine isn't what folks have in mind when they talk about the distribution of copyrighted content.

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#39
post #31
post #27

Earlier quoted context omitted.

No they aren't. Where's the Terms of Service for a book? What about secondhand books or books you get from the library or books you borrow from a friend? I never had a book tell me to an accept a license agreement before I could read it.

Are you trolling? Look at the ISBN/info page of literally any book. It will say something like 'all rights reserved' and that you can't reproduce any part without permission, except for limited cases. Please go get a book off of your shelf and look, I'm begging you. Edit - Example from Infinite Jest: https://burnsiderarebooks.cdn.bibliopolis.com/pictures/14094...

And where does it say "you can't learn anything from this book at all".

We're not talking about copying verbatim, that's the whole point.

Re: OpenAI now tries to hide that ChatGPT was trained on copyrighted books

#40
post #25
post #20

Earlier quoted context omitted.

So aspiring authors should avoid all reading, because they need licenses for all works that inspired them?

That is an absurd conclusion to draw. Authors do not need licenses for works that inspired them.

Then why should AI?
Post reply on HN