Live data from Hacker News

Flux: Open-source text-to-image model with 12B parameters

blog.fal.ai

191–200 of 239 posts

Re: Flux: Open-source text-to-image model with 12B parameters

#191
post #110

Earlier quoted context omitted.

The playground is a drag. After accepting being forced to sign up, attach my GitHub, and hand over my email address, I entered the desired prompt and waited with anticipation.. Only to see a black screen and how much it's going to cost per megapixel. Bummer. After seeing what was generated in the blog post I was excited to try it! Now feeling disappointed. I was hoping it'd be more like https://play.go.dev . Good luc…

https://replicate.com/black-forest-labs/flux-dev is working very nicely. No sign-up.

I was inspired by the SD3 problems to use this prompt:

"a woman lying on her back wearing a blouse and shorts."

But it wouldn't render the image - i instead got a NSFW warning. That's one way to hide the fact that it cannot render it properly i guess...

PS: after a few tries it rendered "a woman lying on her back" correctly.

Re: Flux: Open-source text-to-image model with 12B parameters

#192
post #179
post #142

Earlier quoted context omitted.

I'm personally comfortable calling a model "open source" if the license is compatible with the https://opensource.org/ definition. The Llama models aren't. Some of the Mistral models are (the Apache 2 ones). Microsoft Phi-3 is - it's MIT.

Open source must include source material so that another can reproduce that the model. I would expect that to be a minimum.

I agree, but that can't happen with the vast majority of these models because they're trained on unlicensed data so they can't slap an open source license on the training data and distribute it.

I've decided to draw my personal line at Open Source Initiative compliance for the license they release the model itself under.

I respect the opinion that it's not truly open source unless they release the training data as well, but I've decided not to make that part of my own personal litmus test here.

My reasoning is that knowing something is "open source" helps me decide what I legally can or cannot do with it when building my own software. Not having access to the training data downs affect my legal rights, it just affects my ability to recompile myself. And I don't have millions of dollars of GPUs so that isn't so important to me, personally.

Re: Flux: Open-source text-to-image model with 12B parameters

#193

Earlier quoted context omitted.

A personal bugbear is the AI fascination with calling themselves open source, virtue signalling I guess. Open weights is exactly right. Source code and arguably more important datasets are both required to replicate the work, which is more in the spirit of open source (and science). I think Meta is especially egregious here, given their history. Never underestimate the value of getting hordes of unpaid workers to ref…

> virtue signalling I'd prefer "false advertising" - it's more direct and without the culture war baggage.

"Open source" is perceived as a virtue, and their claim is false. Thus false virtue claim. Or... virtue signaling.

Re: Flux: Open-source text-to-image model with 12B parameters

#194
post #178

Earlier quoted context omitted.

My go-to test for these tools so far has been the seven horned, seven eyed lamb mentioned in the Book of Revelation. Every tool I've tried has failed at this task.

Ah. I try the following: > A Gary Larsen, "Far Side" comic of a racoon disguising itself by wearing a fedora and long trench coat. The raccoon's face is mostly hidden by the fedora. There are extra paws sticking out of the front of the trench coat from between the buttons, suggesting that the racoon is in fact a stack of several raccoons. Every human I've ever described this to has no problem picturing what I mean. I…

This gets interesting. One approach that I've used with image generation before is to find an image of the sort that I want, and have Dall-e describe it... and then modify the prompt that it provides to be one with the elements that I want.

The first attempt at this based on https://reductress.com/post/my-boyfriends-are-always-two-kid... ... really misunderstood the image. This may also be part of the problem.

    The image shows an imaginative, whimsical illustration of a character composed of two parts. The upper part features a man dressed in a long, elegant gray coat, wearing a bowler hat and round sunglasses, with a sophisticated white polka-dot ascot tie. His face has a subtle smile. The lower part of the character transitions seamlessly into a smaller figure of a cat, appearing to wear striped pants, with its tail visible. The entire character combines human and feline elements, creating a surreal, anthropomorphic appearance. The illustration is in black and white, emphasizing a stylized, cartoon-like design.
I then went to the image from https://www.reddit.com/r/DnD/comments/c6fdw4/oc_introducing_...

And that provided:

    The image captures a whimsical and secretive scene featuring three dwarves stacked in a totem formation, each attempting to conceal their nature under a large brown cloak. The top dwarf has a bright, cheerful expression and blond hair, holding the cloak wide to mimic wings, and is dressed in black armor adorned with teal gems and matching earrings. The middle dwarf displays a fierce expression, sporting a bushy orange beard, and is also clad in similar dark armor with teal embellishments. The bottom dwarf, an older figure with a long white beard, is adorned in a royal dark outfit with gold accents and a small crown, clasping a glowing white orb. This trio of dwarves, each with distinctive fantasy armor, unites in a playful attempt to disguise their stature and nature, adding an element of adventure and mystery to the scene.
Working off of that idea of the totem formation ... "Create an image featuring three children in a totem pole formation that are trying to conceal their nature in a single oversized trench coat."

That produced https://imgur.com/a/Of9FsJl

I suspect the orange beard came from the previous part in the session. But that might be an approach to take in trying to describe it in a way that can be used.

Re: Flux: Open-source text-to-image model with 12B parameters

#195
post #89

Earlier quoted context omitted.

> (available without sign-in) FLUX.1 [schnell] (Apache 2.0, open weights, step distilled): https://fal.ai/models/fal-ai/flux/schnell Well, I was wondering about bias in the model, so I entered "a president" as the prompt. Looks like it has a bias alright, but it's even more specific than I expected...

You weren’t kidding. Tried three times and all three were variations of the same[0]. [0] https://fal.media/files/elephant/gu3ZQ46_53BUV6lptexEh.png

[deleted]

Re: Flux: Open-source text-to-image model with 12B parameters

#197
post #15

Earlier quoted context omitted.

Thank you. Their website is super hard to navigate and I can't find a "DOWNLOAD" button.

Note that actually running the model without a A100 GPU or better will be tricker than usual given its size (12B parameters, 24GB on disk). There is a PR to that repo for a diffusers implementation, which may run on a cheap L4 GPU w/ enable_model_cpu_offload(): https://huggingface.co/black-forest-labs/FLUX.1-schnell/comm...

Got it running. But it is a special setup.

* NVIDIA Jetson AGX Orin Dev. Kit with 64 GB shared RAM.

* Default configuration for flux-dev. (FP16, 50 steps)

* 33GB GPU RAM usage.

* 4 minutes 20 seconds per image at around 50 Watt power usage.

Re: Flux: Open-source text-to-image model with 12B parameters

#199
post #194
post #178

Earlier quoted context omitted.

Ah. I try the following: > A Gary Larsen, "Far Side" comic of a racoon disguising itself by wearing a fedora and long trench coat. The raccoon's face is mostly hidden by the fedora. There are extra paws sticking out of the front of the trench coat from between the buttons, suggesting that the racoon is in fact a stack of several raccoons. Every human I've ever described this to has no problem picturing what I mean. I…

This gets interesting. One approach that I've used with image generation before is to find an image of the sort that I want, and have Dall-e describe it... and then modify the prompt that it provides to be one with the elements that I want. The first attempt at this based on https://reductress.com/post/my-boyfriends-are-always-two-kid... ... really misunderstood the image. This may also be part of the problem. The im…

It looks like imgur is blocking Mullvad VPN connections

Re: Flux: Open-source text-to-image model with 12B parameters

#200

I wonder if the key behind the quality of the MidJourney models, and this models, is less about size + architecture and more about the quality of images trained on. It looks like this is the case for LLMs, that the training quality of the data has a significant impact on the output quality of the model, which makes sense. So the real magic is in designing a system to curate that high quality data.

Midjourney unquestionably has heavy data set curation and uses RLHF from users. You don't have to speculate on this as you can see that custom models for SDXL for instance perform vastly better than vanilla SDXL at the same number of parameters. It's all data set and tagging.

SDXL used RLHF too
Post reply on HN