Earlier quoted context omitted.
> I mean, from my perspective, the skill in these (and DALL-E's) image reproductions is truly astonishing. A basic part of it is that neural networks combine learning and memorizing fluidly inside them, and these networks are really really big, so they can memorize stuff good. So when you see it reproduce a Shiba Inu well, don’t think of it as “the model understands Shiba Inus”. Think of it as making a collage out of…
To be clear, I understand the general techniques about (a) how diffusion models can be used to upsample images and generate more photorealistic (or even "cartoon realistic") results and (b) I understand how they can do basic matching of "someone typed in Shiba Inu, look for images of Shiba Inus". What I don't understand is how they do the composition . E.g. for "A giant cobra snake on a farm. The snake is made out of…
Convolutional filters lend themselves to rich combinatorics of compositions[1]: think of them as of context-dependent texture-atoms, repulsing and attracting over the variations of the local multi-dimensional context in the image. The composition is literally a convolutional transformation of local channels encoding related principal components of context.
Astronomical amounts of computations spent via training allow the network to form a lego-set of these texture-atoms in a general distribution of contexts.
At least this is my intuition for the nature of the convnets.
1. https://microscope.openai.com/models/contrastive_16x/image_b...