Earlier quoted context omitted.
To put things in perspective, the dataset it's trained on is ~240TB and Stability has over ~4000 Nvidia A100 (which is much faster than a 1080ti). Without those ingredients, you're highly unlikely to get a model that's worth using (it'll produce mostly useless outputs). That argument also makes little sense when you consider that the model is a couple gigabytes itself, it can't memorize 240TB of data, so it "learned"…
> when you consider that the model is a couple gigabytes itself, it can't memorize 240TB of data, so it "learned". This is just lossy compression with a large and well-tuned (to the expected problem domain) dictionary. Video compression codecs can achieve a 500x compression ratio, and they are general-purpose.
Uncompressed, LAION-5B would be 4PB, for a compression ratio into SD of ~780kx, or one byte per picture.