Tested it using prompts from ideogram (login walled) which has great prompt adherence. Flux generated very very good images. I have been playing with ideogram but i don't want their filters and want to have a similar powerful system running locally. If this runs locally, this is very very close to that in terms of both image quality and prompt adherence. I did fail at writing text clearly when text was a bit complica…
Flux: Open-source text-to-image model with 12B parameters
221–230 of 239 posts
Re: Flux: Open-source text-to-image model with 12B parameters
#222Earlier quoted context omitted.
You're not. I'm surprised at their selections because neither the cooking one nor the beach one adhere to the prompt in very well, and that first one only does because it prompt largely avoids much detail altogether. Overall, the announcement gives the sense that it can make pretty pictures but not very precise ones.
Well, that's nothing new, but it doesn't matter to dedicated users because they don't control it just by typing in text prompts. They use ComfyUI, which is a node editor.
Re: Flux: Open-source text-to-image model with 12B parameters
#223I wonder if the key behind the quality of the MidJourney models, and this models, is less about size + architecture and more about the quality of images trained on. It looks like this is the case for LLMs, that the training quality of the data has a significant impact on the output quality of the model, which makes sense. So the real magic is in designing a system to curate that high quality data.
I would agree - midjourney is getting a free labour since many of their generations are not in secret mode (require pro/mega subscription) so prompts and outputs are visible to everyone. Midjourney rewards users to rating those generations. I wouldn't be surprised if there are some bots on their discord that are scraping those data for training their own models.
Re: Flux: Open-source text-to-image model with 12B parameters
#224I'm really impressed at its ability to output pixel art sprites. Maybe the best general-purpose model I've seen capable of that. In many cases its better than purpose-built models.
Re: Flux: Open-source text-to-image model with 12B parameters
#225Nice one. Will it plan to support both text and image to image?
Re: Flux: Open-source text-to-image model with 12B parameters
#226Great product. BTW I am new to this technology can you please tell me what is the parameter given to Model to make it look like real life image ?
Re: Flux: Open-source text-to-image model with 12B parameters
#227Earlier quoted context omitted.
First they came for fashion magazines, and I said nothing.
Must we always jump to Nazis? This is like the fifth time I see someone paraphrasing Niemöller in an ai context, and it's exhausting. It's also near impossible to take the paraphraser seriously. More to the point, AI is a tool. I could just as well infringe on vanity fair IP using ms-paint. Someone more artistic than me could make a oil-on-canvas copy of their logo too. Or, to turn your own annoying "argument" agains…
As to your use of the argument in the other direction, I’d say it doesn’t work very well because no one with any power is coming for those things.
Re: Flux: Open-source text-to-image model with 12B parameters
#228Earlier quoted context omitted.
First they came for fashion magazines, and I said nothing.
Just to be clear: you're comparing the collapse of the creative restrictions which the state has cleverly branded "intellectual property" to... the holocaust? Of all of the instances on HN of Godwin's law playing out that I've ever seen, this one is the new cake-taker.
Re: Flux: Open-source text-to-image model with 12B parameters
#229Earlier quoted context omitted.
First they came for fashion magazines, and I said nothing.
All journalism is just duplicating the works and performance of others without their permission for profit anyway.
Re: Flux: Open-source text-to-image model with 12B parameters
#230Earlier quoted context omitted.
Well, that's nothing new, but it doesn't matter to dedicated users because they don't control it just by typing in text prompts. They use ComfyUI, which is a node editor.
Does this afford better prompt adherence control in some way?