Live data from Hacker News

OpenAI destroyed a trove of books used to train AI models

businessinsider.com

31–38 of 38 posts

Re: OpenAI destroyed a trove of books used to train AI models

#31

Earlier quoted context omitted.

Yup. The big pretense we pull as an industry is to pretend all of the data for all these models are somehow legitimate. It's all illegal. But what are you gonna do about it?

I think anyone who wants to opt out of being in the training data for LLMs should be able to just like anyone who doesn’t want their website indexed by Google should also be able to opt out.

[dead]

Re: OpenAI destroyed a trove of books used to train AI models

#32
post #25

Earlier quoted context omitted.

If something is a legal liability to have, non-use is probably 2 seconds after you finish using it (in their case, finishing training, just keeping the weights)

Exactly. The reason that they stopped using the dataset and the reason they deleted it are likely the same - legal liability. Obviously not in their interests to state that but when this is the best alternative explanation they can offer they might as well have. Personally I support the use of these books for training AI but I think this needs to be decided in court and/or with legislation, not hidden under the carpe…

>Personally I support the use of these books for training AI

I need a better understanding of postgres. Can you write me an A-Z book on postgres? I won't actually buy it from you, just grab the text, train a model, and get the model to answer all questions I have. But the book would be super helpful please. Oh I'm also going to sell a service on this model too...because I like money.

I get sarcasm isn't exactly a great form of debate, but it felt suitable here. I ABSOLUTELY understand why people don't want AI training on their books.

Re: OpenAI destroyed a trove of books used to train AI models

#33

Earlier quoted context omitted.

Exactly. The reason that they stopped using the dataset and the reason they deleted it are likely the same - legal liability. Obviously not in their interests to state that but when this is the best alternative explanation they can offer they might as well have. Personally I support the use of these books for training AI but I think this needs to be decided in court and/or with legislation, not hidden under the carpe…

>Personally I support the use of these books for training AI I need a better understanding of postgres. Can you write me an A-Z book on postgres? I won't actually buy it from you, just grab the text, train a model, and get the model to answer all questions I have. But the book would be super helpful please. Oh I'm also going to sell a service on this model too...because I like money. I get sarcasm isn't exactly a gre…

Thanks, that's a very interesting perspective which I hadn't previously considered.

Re: OpenAI destroyed a trove of books used to train AI models

#34
post #29
post #28

Earlier quoted context omitted.

Not really no, I've been in that situation and we were all aware that we were committing piracy in the company's name.

Were you committing a criminal offence or tort, and was it in your personal or corporate capacity?

Well, since we were asked to locate material to pirate, which the company intended to sell, it would be criminal, but the criminal act would be the sale of the material, not the procurement of the material, as I read the law.

There might have been a possibility of a claim that we employees asked to download the material were committing a criminal offense under the 'posession with during business with intent to commit and act of infringement', but even there I think it'd be the VPs that had the stupid idea that were committing infringement, not us lowly employees asked to acquire the material.

About the best I suspect CPS could hope for would be a threat as a cudgel to admit who suggested the idea. But I'd have happily told them, and I left the company very shortly after.

Re: OpenAI destroyed a trove of books used to train AI models

#35
post #21

Earlier quoted context omitted.

But is training a model "selling derivatives"? And is so, of what?

If it so, that's copyright infringement. The pending litigations are, in part, to resolve the dispute about whether or not they are "selling derivates." I'm not sure why you conflate your own personal lack of knowledge about these matters with a good argument against the copyright holders.

Judging from your derogatory replies, you hold a strong opinion on the matter.

Is that opinion based on facts and actual cases? Because as far as I know, the entire reason we have these law suits as presented in TLA, is to find out how the law stands on the exact questions that I pose.

So far, it seems copyright (in its current form) isn't suitable for (content) creators to prohibit LLM-researchers and -providers from training models on their creations.

Re: OpenAI destroyed a trove of books used to train AI models

#38
post #35

Earlier quoted context omitted.

If it so, that's copyright infringement. The pending litigations are, in part, to resolve the dispute about whether or not they are "selling derivates." I'm not sure why you conflate your own personal lack of knowledge about these matters with a good argument against the copyright holders.

Judging from your derogatory replies, you hold a strong opinion on the matter. Is that opinion based on facts and actual cases? Because as far as I know, the entire reason we have these law suits as presented in TLA, is to find out how the law stands on the exact questions that I pose. So far, it seems copyright (in its current form) isn't suitable for (content) creators to prohibit LLM-researchers and -providers fro…

Which opinion?
Post reply on HN