Earlier quoted context omitted.
Given that this seems to happen all the time without antitrust issues it probably wouldn't, even though I feel like it should. What we need is a legal way for companies to keep the data open, but also require OpenAI and friends to pay them for it.
> What we need is a legal way for companies to keep the data open, but also require OpenAI and friends to pay them for it. Couldn't that be accomplished by a law or ruling that using something for training AI doesn't exempt you from having to follow its license? OpenAI is already in blatant violation of both the "BY" and "SA" parts of the existing license.
Let's say I take a collection of images and use a program to compress them. When decompressed, the images are close to, but not exactly the same as the originals. Despite being in a different format, and despite not being exactly the same as the originals, the copyright to the compressed images is still held by whoever previously held it.
If I take the collection of images from earlier and train a diffusion model based on it, I'm essentially just compressing it a different way. With the right prompt, you can get out something very similar to what you put in.