Earlier quoted context omitted.
> If you have to already have access to the copyrighted images to find them in the model, the argument seems weak. That makes no sense. The copyright holder has access to their own inventions, of course. That's the standard in any copyright claim. > A sufficiently advanced model could, in theory, generate any image. You could then, again in theory, find an embedding for any image. Does said model then infringe on all…
> That makes no sense. The copyright holder has access to their own inventions, of course. No, not talking about the copyright holder, I'm talking about the hypothetical individual(s) creating infringing copies. If those people need to already have a copy of the image to extract a copy of the image from the generative image model, then I'm saying the argument that the model itself is infringing seems weak. Or it's at…
AFAIK, they don't need that to extract it.
> Instead, you already have to have a copy of an image to find an embedding.
You literally always need a copy or a suitable imprecise hash of the original to test for infringement. How else would you know which copyright was infringed? But it's not a matter of a 1-1 match (see below).
> However, it's a case-by-case thing.
Of course, it's a case-by-case thing, as the OP also indicated. It's easy to show mathematically that these models cannot compress well enough to contain all training images. The question is how much they infringe on some of them.
Bear in mind that it is not necessary at all to create a perfect copy of an image the infringe copyright. As I've stated above, even a lousy and mostly incorrect rendition of a pop song in a street cafe may infringe copyright. The makers behind the song "Blurred Lines" lost a lawsuit because the cowbell rhythm in the background was similar to that of another song. That and the "feeling" was similar.
The same is true for images. What counts are criteria like artistic originality, intent, subjective similarities, experts laying out similarities in style, and so on.
I mean, don't get me wrong, I understand perfectly well what you're trying to argue for. All I'm saying is that it doesn't match the reality of how the law deals with copyright.