Isn't copyright about reproducing and/or distributing? Does training a model count as reproduction or distribution? Or can copyright say something about how I consume my books? Can copyright prohibit me to light a fire, or whipe my behind with, say Harry Potter and the goblet of fire?
Well you usually cannot sell derivative works as your own.
OpenAI destroyed a trove of books used to train AI models
21–30 of 38 posts
Re: OpenAI destroyed a trove of books used to train AI models
#22Problem statement for poor person’s pre training: Where to get lots of data for use if they aren’t providing it as data sets and we can’t share them? And for multimodal models? And less risk in copyright?
The idea was to buy lots of encyclopedia sets, school curriculums, picture books… used media at low prices (esp bin sales) full of information. Digitize them with book scanners. Keep the digital copies and throw away the physical copies. Now, you have a huge, training set of legal data acquired dirt cheap with preprocessing allowing cheaper labor.
From there, use the copies in places like Japan where it’s legal to use them for AI training so long as one has legal access. This also incentivizes the owner to pre-train the model themselves so there no distribution of original works. Also, I envisioned people partly training models with their data, handing the weights off to another company, they add theirs to it, and so on. Daisy chain the process with only the models, not the copyrighted works, distributed. My copyright amendment added preprocessing and copying for just this purpose to avoid these ridiculous hacks.
To be clear, I wouldn’t do this without consulting several lawyers on what was clear. I’d rather not be destroying books and filling landfills. It is pure speculation I made due to how overly-strong copyright is in my country. It assumes we can use copyrighted works (a) for personal use and (b) with a physical to digital conversion. If not, I also thought those would be easier rights to fight for, too.
However, legal hacks like scanning used works to be trained in other countries might be all we have if legal systems don’t adapt to what society is currently doing with copyrighted works. I mean in a way that’s fair to all sides rather than benefiting only one. I’m up for compromise. Pro-copyright side usually hasn’t been, though.
Re: OpenAI destroyed a trove of books used to train AI models
#23Isn't copyright about reproducing and/or distributing? Does training a model count as reproduction or distribution? Or can copyright say something about how I consume my books? Can copyright prohibit me to light a fire, or whipe my behind with, say Harry Potter and the goblet of fire?
It cannot.
This whole kerfuffle isn't about OpenAI buying a bunch of books and disposing of them in a non-copyright friendly way though.
Re: OpenAI destroyed a trove of books used to train AI models
#24Earlier quoted context omitted.
Well you usually cannot sell derivative works as your own.
But is training a model "selling derivatives"? And is so, of what?
Re: OpenAI destroyed a trove of books used to train AI models
#25"These datasets, created by former employees who are no longer with OpenAI, were last used in 2021 and deleted due to non-use in 2022." - Deleting the dataset because of non-use sounds completely implausible. It says the dataset is 67B tokens, which is less than 1TB of data. Why would you bother to delete it given it would cost more or less nothing to keep?
Re: OpenAI destroyed a trove of books used to train AI models
#26Isn't all/most their training data copyrighted anyways? We just have to say it's fair use, because it is useful to everyone. Maybe just require them to open their model.
Yup. The big pretense we pull as an industry is to pretend all of the data for all these models are somehow legitimate. It's all illegal. But what are you gonna do about it?
Re: OpenAI destroyed a trove of books used to train AI models
#27Re: OpenAI destroyed a trove of books used to train AI models
#28Earlier quoted context omitted.
Is "I was just following orders" a viable defense here?
It might be. Is it reasonable to assume the material was cleared by leagal department if your manager came and told you to download books from this long list and clean them up for training purposes?
Re: OpenAI destroyed a trove of books used to train AI models
#29Earlier quoted context omitted.
It might be. Is it reasonable to assume the material was cleared by leagal department if your manager came and told you to download books from this long list and clean them up for training purposes?
Not really no, I've been in that situation and we were all aware that we were committing piracy in the company's name.
Re: OpenAI destroyed a trove of books used to train AI models
#30"These datasets, created by former employees who are no longer with OpenAI, were last used in 2021 and deleted due to non-use in 2022." - Deleting the dataset because of non-use sounds completely implausible. It says the dataset is 67B tokens, which is less than 1TB of data. Why would you bother to delete it given it would cost more or less nothing to keep?
If something is a legal liability to have, non-use is probably 2 seconds after you finish using it (in their case, finishing training, just keeping the weights)
Obviously not in their interests to state that but when this is the best alternative explanation they can offer they might as well have.
Personally I support the use of these books for training AI but I think this needs to be decided in court and/or with legislation, not hidden under the carpet.