Am I missing something? The beach image they give still fails to follow the prompt in major ways.
Flux: Open-source text-to-image model with 12B parameters
21–30 of 239 posts
Re: Flux: Open-source text-to-image model with 12B parameters
#22I wonder if the key behind the quality of the MidJourney models, and this models, is less about size + architecture and more about the quality of images trained on. It looks like this is the case for LLMs, that the training quality of the data has a significant impact on the output quality of the model, which makes sense. So the real magic is in designing a system to curate that high quality data.
You can get better results with better data, for sure. And better architecture, for sure. But raw size is really important the difference in quality for models, all else held equal, is HUGE and obvious if you play with them.
Re: Flux: Open-source text-to-image model with 12B parameters
#23Re: Flux: Open-source text-to-image model with 12B parameters
#24I wonder if the key behind the quality of the MidJourney models, and this models, is less about size + architecture and more about the quality of images trained on. It looks like this is the case for LLMs, that the training quality of the data has a significant impact on the output quality of the model, which makes sense. So the real magic is in designing a system to curate that high quality data.
Re: Flux: Open-source text-to-image model with 12B parameters
#25Re: Flux: Open-source text-to-image model with 12B parameters
#26Earlier quoted context omitted.
Note that actually running the model without a A100 GPU or better will be tricker than usual given its size (12B parameters, 24GB on disk). There is a PR to that repo for a diffusers implementation, which may run on a cheap L4 GPU w/ enable_model_cpu_offload(): https://huggingface.co/black-forest-labs/FLUX.1-schnell/comm...
3090 TIs should be able to handle it without much in the way of tricks for a "reasonable" (for the HN crowd) price.
Re: Flux: Open-source text-to-image model with 12B parameters
#27Looks like a very promising model. Hope to see the comfyui community get it going quickly
Re: Flux: Open-source text-to-image model with 12B parameters
#28(available without sign-in) FLUX.1 [schnell] (Apache 2.0, open weights, step distilled): https://fal.ai/models/fal-ai/flux/schnell
(requires sign-in) FLUX.1 [dev] (non-commercial, open weights, guidance distilled): https://fal.ai/models/fal-ai/flux/dev
FLUX.1 [pro] (closed source [only available thru APIs], SOTA, raw): https://fal.ai/models/fal-ai/flux-pro
Re: Flux: Open-source text-to-image model with 12B parameters
#29Photo of teen girl in a ski mask making an origami swan in a barn. There is caption on the bottom of the image: "EAT DRUGS" in yellow font. In the background there is a framed photo of obama
https://i.imgur.com/RifcWZc.png
Donald Trump on the cover of "Leopards Ate My Face" magazine
Re: Flux: Open-source text-to-image model with 12B parameters
#30If this runs locally, this is very very close to that in terms of both image quality and prompt adherence.
I did fail at writing text clearly when text was a bit complicated. This ideogram image's prompt for example https://ideogram.ai/g/GUw6Vo-tQ8eRWp9x2HONdA/0
> A captivating and artistic illustration of four distinct creative quarters, each representing a unique aspect of creativity. In the top left, a writer with a quill and inkpot is depicted, showcasing their struggle with the text "THE STRUGGLE IS NOT REAL 1: WRITER". The scene is comically portrayed, highlighting the writer's creative challenges. In the top right, a figure labeled "THE STRUGGLE IS NOT REAL 2: COPY ||PASTER" is accompanied by a humorous comic drawing that satirically demonstrates their approach. In the bottom left, "THE STRUGGLE IS NOT REAL 3: THE RETRIER" features a character retrieving items, complete with an entertaining comic illustration. Lastly, in the bottom right, a remixer, identified as "THE STRUGGLE IS NOT REAL 4: THE REMI
Otherwise, the quality is great. I stopped using stable diffusion long time ago, the tools and tech around it became very messy, its not fun anymore. Been using ideogram for fun but I want something like ideogram that I can run locally without any filters. This is looking perfect so far.
This is not ideogram, but its very very good.