Live data from Hacker News

OpenAI destroyed a trove of books used to train AI models

businessinsider.com

21–30 of 38 posts

Re: OpenAI destroyed a trove of books used to train AI models

#21
post #5

Isn't copyright about reproducing and/or distributing? Does training a model count as reproduction or distribution? Or can copyright say something about how I consume my books? Can copyright prohibit me to light a fire, or whipe my behind with, say Harry Potter and the goblet of fire?

Well you usually cannot sell derivative works as your own.

But is training a model "selling derivatives"? And is so, of what?

Re: OpenAI destroyed a trove of books used to train AI models

#22
There was one scenario like this I came up with when brainstorming copyright issues. In the past, articles I read said you could have a backup copy of a copyrighted work that stayed with you. Moving one in a different direction than the other might be a violation since it could be distribution of copies. But, we might be able to make one, digital copy of a physical work for our own use. How to use that?

Problem statement for poor person’s pre training: Where to get lots of data for use if they aren’t providing it as data sets and we can’t share them? And for multimodal models? And less risk in copyright?

The idea was to buy lots of encyclopedia sets, school curriculums, picture books… used media at low prices (esp bin sales) full of information. Digitize them with book scanners. Keep the digital copies and throw away the physical copies. Now, you have a huge, training set of legal data acquired dirt cheap with preprocessing allowing cheaper labor.

From there, use the copies in places like Japan where it’s legal to use them for AI training so long as one has legal access. This also incentivizes the owner to pre-train the model themselves so there no distribution of original works. Also, I envisioned people partly training models with their data, handing the weights off to another company, they add theirs to it, and so on. Daisy chain the process with only the models, not the copyrighted works, distributed. My copyright amendment added preprocessing and copying for just this purpose to avoid these ridiculous hacks.

To be clear, I wouldn’t do this without consulting several lawyers on what was clear. I’d rather not be destroying books and filling landfills. It is pure speculation I made due to how overly-strong copyright is in my country. It assumes we can use copyrighted works (a) for personal use and (b) with a physical to digital conversion. If not, I also thought those would be easier rights to fight for, too.

However, legal hacks like scanning used works to be trained in other countries might be all we have if legal systems don’t adapt to what society is currently doing with copyrighted works. I mean in a way that’s fair to all sides rather than benefiting only one. I’m up for compromise. Pro-copyright side usually hasn’t been, though.

Re: OpenAI destroyed a trove of books used to train AI models

#23
post #5

Isn't copyright about reproducing and/or distributing? Does training a model count as reproduction or distribution? Or can copyright say something about how I consume my books? Can copyright prohibit me to light a fire, or whipe my behind with, say Harry Potter and the goblet of fire?

> Or can copyright say something about how I consume my books?

It cannot.

This whole kerfuffle isn't about OpenAI buying a bunch of books and disposing of them in a non-copyright friendly way though.

Re: OpenAI destroyed a trove of books used to train AI models

#24
post #21

Earlier quoted context omitted.

Well you usually cannot sell derivative works as your own.

But is training a model "selling derivatives"? And is so, of what?

If it so, that's copyright infringement. The pending litigations are, in part, to resolve the dispute about whether or not they are "selling derivates." I'm not sure why you conflate your own personal lack of knowledge about these matters with a good argument against the copyright holders.

Re: OpenAI destroyed a trove of books used to train AI models

#25

"These datasets, created by former employees who are no longer with OpenAI, were last used in 2021 and deleted due to non-use in 2022." - Deleting the dataset because of non-use sounds completely implausible. It says the dataset is 67B tokens, which is less than 1TB of data. Why would you bother to delete it given it would cost more or less nothing to keep?

If something is a legal liability to have, non-use is probably 2 seconds after you finish using it (in their case, finishing training, just keeping the weights)

Re: OpenAI destroyed a trove of books used to train AI models

#26

Isn't all/most their training data copyrighted anyways? We just have to say it's fair use, because it is useful to everyone. Maybe just require them to open their model.

Yup. The big pretense we pull as an industry is to pretend all of the data for all these models are somehow legitimate. It's all illegal. But what are you gonna do about it?

I think anyone who wants to opt out of being in the training data for LLMs should be able to just like anyone who doesn’t want their website indexed by Google should also be able to opt out.

Re: OpenAI destroyed a trove of books used to train AI models

#27
So when is the author’s guild going to realize that you can pretrain a base model on public domain material and then once it’s distributed anyone can fine tune it on whatever books they want to have their own version that writes in the style of X? On commodity GPUs no less. In 5 years when the current gen GPUs are cheap and there are hundreds of one click fine tuning apps this will seem like an absurd waste of time.

Re: OpenAI destroyed a trove of books used to train AI models

#28

Earlier quoted context omitted.

Is "I was just following orders" a viable defense here?

It might be. Is it reasonable to assume the material was cleared by leagal department if your manager came and told you to download books from this long list and clean them up for training purposes?

Not really no, I've been in that situation and we were all aware that we were committing piracy in the company's name.

Re: OpenAI destroyed a trove of books used to train AI models

#29
post #28

Earlier quoted context omitted.

It might be. Is it reasonable to assume the material was cleared by leagal department if your manager came and told you to download books from this long list and clean them up for training purposes?

Not really no, I've been in that situation and we were all aware that we were committing piracy in the company's name.

Were you committing a criminal offence or tort, and was it in your personal or corporate capacity?

Re: OpenAI destroyed a trove of books used to train AI models

#30
post #25

"These datasets, created by former employees who are no longer with OpenAI, were last used in 2021 and deleted due to non-use in 2022." - Deleting the dataset because of non-use sounds completely implausible. It says the dataset is 67B tokens, which is less than 1TB of data. Why would you bother to delete it given it would cost more or less nothing to keep?

If something is a legal liability to have, non-use is probably 2 seconds after you finish using it (in their case, finishing training, just keeping the weights)

Exactly. The reason that they stopped using the dataset and the reason they deleted it are likely the same - legal liability.

Obviously not in their interests to state that but when this is the best alternative explanation they can offer they might as well have.

Personally I support the use of these books for training AI but I think this needs to be decided in court and/or with legislation, not hidden under the carpet.

Post reply on HN