Feels weird that the Authors Guild so much wants to make public the names of the former OpenAI employees who created the data sets. That seems entirely unnecessary and irrelevant to the case. If what they did was part of their sanctioned work for OpenAI, they are not the ones responsible for how it was used.
OpenAI destroyed a trove of books used to train AI models
11–20 of 38 posts
Re: OpenAI destroyed a trove of books used to train AI models
#12Deleting the dataset because of non-use sounds completely implausible. It says the dataset is 67B tokens, which is less than 1TB of data. Why would you bother to delete it given it would cost more or less nothing to keep?
Re: OpenAI destroyed a trove of books used to train AI models
#13Assuming they train newer models using output from older versions, isn't that data they collected from "books1" and "books2" still encoded in the weights of their current models? This also begs the question, does OpenAI really honor their privacy controls and not use user information to train their models if the user opts out? It seems most companies are operating in "ask for forgiveness than permission"-mode as they…
Re: OpenAI destroyed a trove of books used to train AI models
#14Isn't copyright about reproducing and/or distributing? Does training a model count as reproduction or distribution? Or can copyright say something about how I consume my books? Can copyright prohibit me to light a fire, or whipe my behind with, say Harry Potter and the goblet of fire?
Re: OpenAI destroyed a trove of books used to train AI models
#15That sounds like ass-covering, and maybe destruction of evidence. If the data is destroyed, won't it be much harder to prove which books they violated copyright on and to figure out the damages owed?
Re: OpenAI destroyed a trove of books used to train AI models
#16Assuming they train newer models using output from older versions, isn't that data they collected from "books1" and "books2" still encoded in the weights of their current models? This also begs the question, does OpenAI really honor their privacy controls and not use user information to train their models if the user opts out? It seems most companies are operating in "ask for forgiveness than permission"-mode as they…
Sure, in much the same way if you save a JPEG at 75% quality the data in the image is still encoded. But if you repeat a lossy encoding over and over without saving the original, well... wikipedia has a nice visualization of what happens: https://en.wikipedia.org/wiki/Generation_loss
Re: OpenAI destroyed a trove of books used to train AI models
#17Feels weird that the Authors Guild so much wants to make public the names of the former OpenAI employees who created the data sets. That seems entirely unnecessary and irrelevant to the case. If what they did was part of their sanctioned work for OpenAI, they are not the ones responsible for how it was used.
Re: OpenAI destroyed a trove of books used to train AI models
#18Assuming they train newer models using output from older versions, isn't that data they collected from "books1" and "books2" still encoded in the weights of their current models? This also begs the question, does OpenAI really honor their privacy controls and not use user information to train their models if the user opts out? It seems most companies are operating in "ask for forgiveness than permission"-mode as they…
Re: OpenAI destroyed a trove of books used to train AI models
#19Isn't all/most their training data copyrighted anyways? We just have to say it's fair use, because it is useful to everyone. Maybe just require them to open their model.
Yup. The big pretense we pull as an industry is to pretend all of the data for all these models are somehow legitimate. It's all illegal. But what are you gonna do about it?
I feel the tech industry took the proverb "better ask for forgiveness than permission," then dropped the "forgiveness" part.
Re: OpenAI destroyed a trove of books used to train AI models
#20Feels weird that the Authors Guild so much wants to make public the names of the former OpenAI employees who created the data sets. That seems entirely unnecessary and irrelevant to the case. If what they did was part of their sanctioned work for OpenAI, they are not the ones responsible for how it was used.
Is "I was just following orders" a viable defense here?