Live data from Hacker News

Flux: Open-source text-to-image model with 12B parameters

blog.fal.ai

1–10 of 239 posts

Re: Flux: Open-source text-to-image model with 12B parameters

#4
You can try the models on replicate https://replicate.com/black-forest-labs.

Result (distilled schnell model) for

"Photo of Criminal in a ski mask making a phone call in front of a store. There is caption on the bottom of the image: "It's time to Counter the Strike...". There is a red arrow pointing towards the caption. The red arrow is from a Red circle which has an image of Halo Master Chief in it."

https://www.reddit.com/r/StableDiffusion/s/SsPeQRJIkw

Re: Flux: Open-source text-to-image model with 12B parameters

#6

You can try the models on replicate https://replicate.com/black-forest-labs . Result (distilled schnell model) for "Photo of Criminal in a ski mask making a phone call in front of a store. There is caption on the bottom of the image: "It's time to Counter the Strike...". There is a red arrow pointing towards the caption. The red arrow is from a Red circle which has an image of Halo Master Chief in it." https://www.re…

[deleted]

Re: Flux: Open-source text-to-image model with 12B parameters

#8
post #5

[flagged]

Hmm I was able to try the model myself (generated a handful of images) and didn't log in.

It’s a fal login, which is linked/federated with GitHub if you have done that before.

It takes zero time to “make an account”.

Re: Flux: Open-source text-to-image model with 12B parameters

#9
Wow.

I have seen a lot of promises made by diffusion models.

This is in a whole different world. I legitimately feel bad for the people still a StabilityAI.

The playground testing is really something else!

The licensing model isn’t bad, although I would like to see them promise to open up their old closed source models under Apache when they release new API versions.

The prompt adherence and the breadth of topics it seems to know without a finetune and without any LORAs, is really amazing.

Re: Flux: Open-source text-to-image model with 12B parameters

#10
The [schnell] model variant is Apache-licensed and is open sourced on Hugging Face: https://huggingface.co/black-forest-labs/FLUX.1-schnell

It is very fast and very good at rendering text, and appears to have a text encoder such that the model can handle both text and positioning much better: https://x.com/minimaxir/status/1819041076872908894

A fun consequence of better text rendering is that it means text watermarks from its training data appear more clearly: https://x.com/minimaxir/status/1819045012166127921

Post reply on HN