Live data from Hacker News

OpenAI destroyed a trove of books used to train AI models

businessinsider.com

11–20 of 38 posts

Re: OpenAI destroyed a trove of books used to train AI models

#11
post #4

Feels weird that the Authors Guild so much wants to make public the names of the former OpenAI employees who created the data sets. That seems entirely unnecessary and irrelevant to the case. If what they did was part of their sanctioned work for OpenAI, they are not the ones responsible for how it was used.

Until OpenAI takes the blame, why would the Author's Guild not pursue their claim?

Re: OpenAI destroyed a trove of books used to train AI models

#12
"These datasets, created by former employees who are no longer with OpenAI, were last used in 2021 and deleted due to non-use in 2022." -

Deleting the dataset because of non-use sounds completely implausible. It says the dataset is 67B tokens, which is less than 1TB of data. Why would you bother to delete it given it would cost more or less nothing to keep?

Re: OpenAI destroyed a trove of books used to train AI models

#13
post #6

Assuming they train newer models using output from older versions, isn't that data they collected from "books1" and "books2" still encoded in the weights of their current models? This also begs the question, does OpenAI really honor their privacy controls and not use user information to train their models if the user opts out? It seems most companies are operating in "ask for forgiveness than permission"-mode as they…

[deleted]

Re: OpenAI destroyed a trove of books used to train AI models

#14
post #5

Isn't copyright about reproducing and/or distributing? Does training a model count as reproduction or distribution? Or can copyright say something about how I consume my books? Can copyright prohibit me to light a fire, or whipe my behind with, say Harry Potter and the goblet of fire?

Questions like these are so far below a level of reasonable discourse that they are detrimental.

Re: OpenAI destroyed a trove of books used to train AI models

#15
> The unsealed letter from OpenAI's lawyers, which is labeled "highly confidential - attorneys' eyes only," says that the use of "books1" and "books2" for model training was discontinued in late 2021 and that the datasets were deleted in mid-2022 because of their nonuse. The letter goes on to say that none of the other data used to train GPT-3 has been deleted and offers attorneys for the Authors Guild access to those other datasets.

That sounds like ass-covering, and maybe destruction of evidence. If the data is destroyed, won't it be much harder to prove which books they violated copyright on and to figure out the damages owed?

Re: OpenAI destroyed a trove of books used to train AI models

#16
post #6

Assuming they train newer models using output from older versions, isn't that data they collected from "books1" and "books2" still encoded in the weights of their current models? This also begs the question, does OpenAI really honor their privacy controls and not use user information to train their models if the user opts out? It seems most companies are operating in "ask for forgiveness than permission"-mode as they…

> Assuming they train newer models using output from older versions, isn't that data they collected from "books1" and "books2" still encoded in the weights of their current models?

Sure, in much the same way if you save a JPEG at 75% quality the data in the image is still encoded. But if you repeat a lossy encoding over and over without saving the original, well... wikipedia has a nice visualization of what happens: https://en.wikipedia.org/wiki/Generation_loss

Re: OpenAI destroyed a trove of books used to train AI models

#17
post #4

Feels weird that the Authors Guild so much wants to make public the names of the former OpenAI employees who created the data sets. That seems entirely unnecessary and irrelevant to the case. If what they did was part of their sanctioned work for OpenAI, they are not the ones responsible for how it was used.

Is "I was just following orders" a viable defense here?

Re: OpenAI destroyed a trove of books used to train AI models

#18
post #6

Assuming they train newer models using output from older versions, isn't that data they collected from "books1" and "books2" still encoded in the weights of their current models? This also begs the question, does OpenAI really honor their privacy controls and not use user information to train their models if the user opts out? It seems most companies are operating in "ask for forgiveness than permission"-mode as they…

They do not train newer models on output from older versions. In fact, they deliberately try to filter out LLM-generated text from their scraped internet datasets.

Re: OpenAI destroyed a trove of books used to train AI models

#19

Isn't all/most their training data copyrighted anyways? We just have to say it's fair use, because it is useful to everyone. Maybe just require them to open their model.

Yup. The big pretense we pull as an industry is to pretend all of the data for all these models are somehow legitimate. It's all illegal. But what are you gonna do about it?

> Yup. The big pretense we pull as an industry is to pretend all of the data for all these models are somehow legitimate. It's all illegal. But what are you gonna do about it?

I feel the tech industry took the proverb "better ask for forgiveness than permission," then dropped the "forgiveness" part.

Re: OpenAI destroyed a trove of books used to train AI models

#20
post #4

Feels weird that the Authors Guild so much wants to make public the names of the former OpenAI employees who created the data sets. That seems entirely unnecessary and irrelevant to the case. If what they did was part of their sanctioned work for OpenAI, they are not the ones responsible for how it was used.

Is "I was just following orders" a viable defense here?

It might be. Is it reasonable to assume the material was cleared by leagal department if your manager came and told you to download books from this long list and clean them up for training purposes?
Post reply on HN