Is this intended to replace Stable Diffusion? Somebody want to give the eli5?
Hope not. This is a worse license.
141–150 of 237 posts
Is this intended to replace Stable Diffusion? Somebody want to give the eli5?
Hope not. This is a worse license.
Is this intended to replace Stable Diffusion? Somebody want to give the eli5?
DeepFloyd IF is a state-of-the-art text-to-image model released on a non-commercial, research-permissible license that provides an opportunity for research labs to examine and experiment with advanced text-to-image generation approaches. In line with other Stability AI models, Stability AI intends to release a DeepFloyd IF model fully open source at a future date.
Is this intended to replace Stable Diffusion? Somebody want to give the eli5?
> Stability AI releases DeepFloyd IF, a powerful text-to-image model Hope not. This is a worse license.
Is this intended to replace Stable Diffusion? Somebody want to give the eli5?
DeepFloyd IF is based on Google's Imagen model, which has two key differences from Stable Diffusion: (1) it denoises in pixel space instead of a compressed latent space, and (2) it uses a 10x larger pretrained text encoder (T5-XXL-1.1) compared to SD's CLIP encoder. (1) allows it to better render high-frequency details and text, and (2) allows it to understand complex prompts much better. These improvements come at the cost of multiple times more memory usage and compute requirements compared to SD, though.
In terms of "will it replace SD?"—in the short term I think yes. But I still think latent diffusion models are the future. For example, Stability is gearing up to release Stable Diffusion XL right now, a larger version of the original SD that does higher fidelity and higher resolution generations. I wouldn't be surprised if it takes the crown back from DeepFloyd when it releases, but I guess we'll have to see.
Is this intended to replace Stable Diffusion? Somebody want to give the eli5?
> Stability AI releases DeepFloyd IF, a powerful text-to-image model Hope not. This is a worse license.
At a later point the model will be renamed "StableIf", and released with a similar license to StableDiffusion.
Is this intended to replace Stable Diffusion? Somebody want to give the eli5?
>DeepFloyd IF works in pixel space. The diffusion is implemented on a pixel level, unlike latent diffusion models (like Stable Diffusion), where latent representations are used.
Wow this does so well on text! The original model struggled a lot, it's impressive to see how far they've come.
Wow this does so well on text! The original model struggled a lot, it's impressive to see how far they've come.