Earlier quoted context omitted.
> Part of my point is that you don't need to produce literally equivalent output. Again, if I record and compress "Revenge of the Sith", there's literally zero pixels shared between my recording and the actual movie. Cool, so I can go upload it for free then right? No, I can't. That's because you would be redistributing the actual material, just in a really roundabout way. GenAI models are not that, they're not a dat…
> That's because you would be redistributing the actual material, just in a really roundabout way. Right, which I’m arguing is what LLMs do just in an even more roundabout way. The technical details of LLMs don’t actually matter. We don’t really care if they’re a database or not. The question is do they reproduce the source material? And yeah, pretty much they do, in a lot of instances. Not all, but a lot. To produce…
Okay, but that is not what's happening here. Demonstratably so. The fact that a model is technically capable of overfitting to certain very repeated points in the training data doesn't mean the entire thing has to be shot down. The non-infringing uses far outweigh the offending ones, by a lot.
If what you say is true, and they do outright copy a lot, then it should be pretty easy for any IP holder to sue anyone who misuses the model that way for copyright infringement on those specific outputs.