Earlier quoted context omitted.
It might not be storing all of the images, but it's clearly storing copies of some of the source images. A few weeks before OpenAI made DALL-E 2 open, I went through a series of prompts using well known artists as a basis. Vermeer was one of them. It generated some pretty amazing works that were 'inspired' but not direct copies of Vermeer's works. I started feeding the same prompts into StableDiffiusion (the HuggingF…
Do you have the exact prompt that essentially generated Vemeer's works? Omitted in the parent, but I also noted that while it is hard to reconcile the idea of human-directed transformative use with ML models, it might be actually easier to verify the models' plagarism---which is not equivalent to copyright infringement but can be a problem by its own---once you have a right prompt as they are decidable.
https://drop.qoid.us/vermeer.png
And, for a slightly different scenario, here's what I got for 'pikachu':
https://drop.qoid.us/pikachu.png
But I think this sort of thing is the exception rather than the rule (similar to Copilot regurgitating the text of software licenses). Girl with a Pearl Earring is so famous that there are countless copies of it on the internet. I don't know what kind of deduplication process was used for Stable Diffusion, but if you consider multiple photos of the same painting, plus images distorted by compression, or cropped, or embedded in other images, plus attempts to redraw it or even parody it, I can easily imagine that the training set contained thousands of images or more that are essentially the same work. And so the model memorized it, albeit at low resolution.
Similarly, it's seen countless different images of Pikachu.
However, this seems to only apply to the very peak of popularity. I tested some other Pokemon: it does a decent job with Charizard, but with Bulbasaur and Squirtle and Jigglypuff its output only somewhat resembles the character design, and for less popular Pokemon there's no resemblance at all.