Live data from Hacker News

Flux: Open-source text-to-image model with 12B parameters

blog.fal.ai

51–60 of 239 posts

Re: Flux: Open-source text-to-image model with 12B parameters

#52

Seems to do pretty poorly with spatial relationships. "An upside down house" -> regular old house "A horse sitting on a dog" -> horse and dog next to eachother "An inverted Lockheed Martin F-22 Raptor" -> yikes https://fal.media/files/koala/zgPYG6SqhD4Y3y_E9MONu.png

a zebra on top of an elephant worked fine for me

Re: Flux: Open-source text-to-image model with 12B parameters

#53
post #29

WILD Photo of teen girl in a ski mask making an origami swan in a barn. There is caption on the bottom of the image: "EAT DRUGS" in yellow font. In the background there is a framed photo of obama https://i.imgur.com/RifcWZc.png Donald Trump on the cover of "Leopards Ate My Face" magazine https://i.imgur.com/6HdBJkr.png

DT coverlines are very authentic. Garbled just like the real thing.

Re: Flux: Open-source text-to-image model with 12B parameters

#56

I wonder if the key behind the quality of the MidJourney models, and this models, is less about size + architecture and more about the quality of images trained on. It looks like this is the case for LLMs, that the training quality of the data has a significant impact on the output quality of the model, which makes sense. So the real magic is in designing a system to curate that high quality data.

I would agree - midjourney is getting a free labour since many of their generations are not in secret mode (require pro/mega subscription) so prompts and outputs are visible to everyone. Midjourney rewards users to rating those generations. I wouldn't be surprised if there are some bots on their discord that are scraping those data for training their own models.

Re: Flux: Open-source text-to-image model with 12B parameters

#57
whenever I see a new model I always see if it can do engineering diagrams (e.g. "two square boxes at a distance of 3.5mm"), still no dice on this one. https://x.com/seveibar/status/1819081632575611279

Would love to see an AI company attack engineering diagrams head on, my current hunch is that they just aren't in the training dataset (I'm very tempted to make a synthetic dataset/benchmark)

Re: Flux: Open-source text-to-image model with 12B parameters

#58
Holy crap this is amazing. I saw an image with a prompt on reddit and didn't believe it was generated imaged. I thought it must be joke that people are sharing non-generated images in the thread.

Reddit message: https://www.reddit.com/r/StableDiffusion/comments/1ehh1hx/an...

Linked image: https://preview.redd.it/dz3djnish2gd1.png?width=1024&format=...

The prompt:

> Photo of Criminal in a ski mask making a phone call in front of a store. There is caption on the bottom of the image: "It's time to Counter the Strike...". There is a red arrow pointing towards the caption. The red arrow is from a Red circle which has an image of Halo Master Chief in it.

Some of the images I generated using schnell model with 8-10 steps using this prompt. https://imgur.com/a/3mM9tKf

Re: Flux: Open-source text-to-image model with 12B parameters

#59

The [schnell] model variant is Apache-licensed and is open sourced on Hugging Face: https://huggingface.co/black-forest-labs/FLUX.1-schnell It is very fast and very good at rendering text, and appears to have a text encoder such that the model can handle both text and positioning much better: https://x.com/minimaxir/status/1819041076872908894 A fun consequence of better text rendering is that it means text watermarks…

That’s not really fair to conclude that the training data contains vanity fair images since the prompt includes “by Vanity Fair”.

I could write “with text that says Shutterstock” in the prompt but that doesn’t necessairly mean the dataset contains that

Re: Flux: Open-source text-to-image model with 12B parameters

#60
post #59

The [schnell] model variant is Apache-licensed and is open sourced on Hugging Face: https://huggingface.co/black-forest-labs/FLUX.1-schnell It is very fast and very good at rendering text, and appears to have a text encoder such that the model can handle both text and positioning much better: https://x.com/minimaxir/status/1819041076872908894 A fun consequence of better text rendering is that it means text watermarks…

That’s not really fair to conclude that the training data contains vanity fair images since the prompt includes “by Vanity Fair”. I could write “with text that says Shutterstock” in the prompt but that doesn’t necessairly mean the dataset contains that

The logo has the same exact copyrighted typography as the real Vanity Fair logo. I've also reproduced the same-copyrighted-typography with other brands with identical composition as copyrighted images. Just asking it "Vanity Fair cover story about Shrek" at a 3:2 ratio gives it a composition identical to a Vanity Fair cover very consistently (subject is in front of logo typography partially obscuring it)

The image linked has a traditional www watermark in the lower-left as well. Even something innocous as a "Super Mario 64" prompt shows a copyright watermark: https://x.com/minimaxir/status/1819093418246631855

Post reply on HN