Earlier quoted context omitted.
> And I seem to recall there are some theoretical lower bounds on even lossy compression. I'm not sure what your math is coming from and it seems trivially wrong. A single black pixel is a very lossy compression of every image on the internet. A picture of the Facebook logo is a slightly-less-lossy compression of every picture on the internet (the Facebook logo shows up on a lot of websites). I would believe that you…
Ok, that's a really low lower bound. I think you'll agree that it would be a bit absurd to threaten legal action against someone for storing a single black pixel. OTOH Someone might be tempted to start a lawsuit if they believe their image is somehow actually stored in a particular data file. For this to be a viable class action lawsuit to pursue, I think you'd have to subscribe to the belief that it's a form of comp…
The black pixel won't get you sued, but the Facebook logo example I used could get you sued. Specifically by Facebook. There is an image (n = 1) that is substantially similar to the output of your compression algorithm.
That is sort of what Getty's lawsuit alleges. Not that every picture is recallable from an LLM, but that several images that are substantially similar to Getty's images are recallable. The same goes with the NYT's lawsuit and OpenAI.