Live data from Hacker News

Flux: Open-source text-to-image model with 12B parameters

blog.fal.ai

231–239 of 239 posts

Re: Flux: Open-source text-to-image model with 12B parameters

#231

Earlier quoted context omitted.

What's the difference between pro and dev? Is the pro one also 12B parameters? Are the example images on the site (the patagonia guy, lego and the beach potato) generated with dev or pro?

I think they are mainly -dev and -schnell. Both models are 12B. -pro is the most powerful and raw, -dev is guidance distilled version of it and -schnell is step distilled version (where you can get pretty good results with 2-8 steps).

what does guidance distilled mean?

something about pro must be better than dev or it wouldn't be made API-only, but what exactly, how does guidance distilling affect pro it and what quality remains in dev?

Re: Flux: Open-source text-to-image model with 12B parameters

#233
post #59

Earlier quoted context omitted.

That’s not really fair to conclude that the training data contains vanity fair images since the prompt includes “by Vanity Fair”. I could write “with text that says Shutterstock” in the prompt but that doesn’t necessairly mean the dataset contains that

The logo has the same exact copyrighted typography as the real Vanity Fair logo. I've also reproduced the same-copyrighted-typography with other brands with identical composition as copyrighted images. Just asking it "Vanity Fair cover story about Shrek" at a 3:2 ratio gives it a composition identical to a Vanity Fair cover very consistently (subject is in front of logo typography partially obscuring it) The image li…

If the training data includes a public blog post which has a screenshot of a vanity fair piece?

It's like GRRM complaining that LLMs can reproduce chunks of text from his books "they fed my novels into it" Oh yeah? It's definitely not all the parts of your book quoted in millions of places online, including several dedicated wiki style sites? That wouldn't be it, right?

Re: Flux: Open-source text-to-image model with 12B parameters

#234
post #145

whenever I see a new model I always see if it can do engineering diagrams (e.g. "two square boxes at a distance of 3.5mm"), still no dice on this one. https://x.com/seveibar/status/1819081632575611279 Would love to see an AI company attack engineering diagrams head on, my current hunch is that they just aren't in the training dataset (I'm very tempted to make a synthetic dataset/benchmark)

It'll probably come suddenly. It has been fascinating to me watching the journey from Stable Diffusion 1 to 3. SD1 was a very crude model, where putting a word in the prompt might or might not add representations of the word to the image. Eg, using the word "hat" somewhere in the prompt might do literally nothing or suddenly there were hats everywhere. The context of the word didn't mean much to SD1. SD2 was more con…

At a rapid clip is a great unintentional pun here.

Re: Flux: Open-source text-to-image model with 12B parameters

#235

Earlier quoted context omitted.

You're not. I'm surprised at their selections because neither the cooking one nor the beach one adhere to the prompt in very well, and that first one only does because it prompt largely avoids much detail altogether. Overall, the announcement gives the sense that it can make pretty pictures but not very precise ones.

Well, that's nothing new, but it doesn't matter to dedicated users because they don't control it just by typing in text prompts. They use ComfyUI, which is a node editor.

I'd say automatic1111 is more popular. Comfy seems like a rat's nest, unreal shader node flashbacks.

Re: Flux: Open-source text-to-image model with 12B parameters

#236

Anyone know why text-to-image models have so many fewer parameters than text models? Are there any large image models (>70b, 400b, etc)?

If a wurd is misspelt then you notis right away.

If a pixel is just slightly the wrong shade of green, nobody really cares.

Re: Flux: Open-source text-to-image model with 12B parameters

#238

Earlier quoted context omitted.

Must we always jump to Nazis? This is like the fifth time I see someone paraphrasing Niemöller in an ai context, and it's exhausting. It's also near impossible to take the paraphraser seriously. More to the point, AI is a tool. I could just as well infringe on vanity fair IP using ms-paint. Someone more artistic than me could make a oil-on-canvas copy of their logo too. Or, to turn your own annoying "argument" agains…

This isn’t paraphrasing, it’s referencing. The reference has become synonymous with saying “this is a slippery slope for X”. As to your use of the argument in the other direction, I’d say it doesn’t work very well because no one with any power is coming for those things.

Nobody at all is "coming for" fashion magazines, but you sure seem to be "coming for" AI. Whether you have any power or not is besides point.

Whether you are paraphrasing or referencing to a famous confessional poem dealing with the Holocaust, the only reasonable interpretation is that you're comparing with the Holocaust. Even if you were unaware of the phrases origins, that's how anyone who does know where it comes from will interpret it. See other comments drawing the same conclusion for reference.

Again. Ai is a tool. It can produce illegal material, just like a pencil can, or a brush with oil and canvas. How are they different? They are not.

Re: Flux: Open-source text-to-image model with 12B parameters

#239
post #226

Great product. BTW I am new to this technology can you please tell me what is the parameter given to Model to make it look like real life image ?

Try something like "Photo of...", "as photography" or "photorealistic". You can even specify the camera model and lens/exposure settings. You can find these in metadata of your phone photos for example.

Thanks for the reply buddy. But I am still not able to understand how camera model and lens/exposure settings can be used to make the photo real.

Lets, say that you took an image of a flower in a garden and ai has also generated an image of the same flower. When we see these pics side by side we find a lot of difference between them. Origin of my question was "how can we minimize this difference ?". How we can tell the machine that the more the magnitude of a certain parameter the more real it is, not sure if camera settings could help in this case.

Post reply on HN