Live data from Hacker News

Flux: Open-source text-to-image model with 12B parameters

blog.fal.ai

161–170 of 239 posts

Re: Flux: Open-source text-to-image model with 12B parameters

#161

Anyone know why text-to-image models have so many fewer parameters than text models? Are there any large image models (>70b, 400b, etc)?

Diffusion is very efficient encoding/decoding. The only reason that diffusion isn't used for text is because text requires discrete outputs.

But see: https://arxiv.org/pdf/2407.15595

Re: Flux: Open-source text-to-image model with 12B parameters

#164
post #152

whenever I see a new model I always see if it can do engineering diagrams (e.g. "two square boxes at a distance of 3.5mm"), still no dice on this one. https://x.com/seveibar/status/1819081632575611279 Would love to see an AI company attack engineering diagrams head on, my current hunch is that they just aren't in the training dataset (I'm very tempted to make a synthetic dataset/benchmark)

> https://fal.media/files/kangaroo/FwO3j7xFIgpIXepqKDj6h.png Prompt: two square boxes at a distance of 3.5mm. Both boxes have the same size, 10cm.

It’s like spelling out wishes for a djinn—better be crystal clear what you’re thinking of…

Re: Flux: Open-source text-to-image model with 12B parameters

#165
post #79

Earlier quoted context omitted.

On my list of AI concerns, whether or not Vanity Fair has it’s copyright infringed does not appear.

First they came for fashion magazines, and I said nothing.

Just to be clear: you're comparing the collapse of the creative restrictions which the state has cleverly branded "intellectual property" to... the holocaust?

Of all of the instances on HN of Godwin's law playing out that I've ever seen, this one is the new cake-taker.

Re: Flux: Open-source text-to-image model with 12B parameters

#166

Earlier quoted context omitted.

You also might want to "clarify" that it is not open source (and neither are any of the other "open source" models). If you want to call it something, try "open weights", although the usage restrictions make even that a HUGE FUCKING STRETCH. Also, everybody should remember that these models are not copyrightable and you should never agree to any license for them...

A personal bugbear is the AI fascination with calling themselves open source, virtue signalling I guess. Open weights is exactly right. Source code and arguably more important datasets are both required to replicate the work, which is more in the spirit of open source (and science). I think Meta is especially egregious here, given their history. Never underestimate the value of getting hordes of unpaid workers to ref…

Indeed. The data is the main "information source" from which the model is trained.

Re: Flux: Open-source text-to-image model with 12B parameters

#167
post #79

Earlier quoted context omitted.

On my list of AI concerns, whether or not Vanity Fair has it’s copyright infringed does not appear.

First they came for fashion magazines, and I said nothing.

Must we always jump to Nazis?

This is like the fifth time I see someone paraphrasing Niemöller in an ai context, and it's exhausting. It's also near impossible to take the paraphraser seriously.

More to the point, AI is a tool. I could just as well infringe on vanity fair IP using ms-paint. Someone more artistic than me could make a oil-on-canvas copy of their logo too.

Or, to turn your own annoying "argument" against you:

First they came for AI models, and I did not speak out, because I wasn't using them. Then they came for Photoshop, and I did not speak out, because I had never learned to use it. Then they came for for oil and canvas, and now there are no art forms left for me.

Re: Flux: Open-source text-to-image model with 12B parameters

#168
post #113

These venture funded startups keep releasing models for free without a business model in sight. I am all for open source but worry it is not sustainable long term.

At this point the only thing an AI startup has to do to get people to spend money on the model is to:

-not censor it

-not be doing prompt injection

It's very easy, which is why no other firm is capable of it.

Re: Flux: Open-source text-to-image model with 12B parameters

#169

hi friends! burkay from fal.ai here. would like to clarify that the model is NOT built by fal. all credit should go to Black Forest Labs ( https://blackforestlabs.ai/ ) which is a new co by the OG stable diffusion team. what we did at fal is take the model and run it on our inference engine optimized to run these kinds of models really really fast. feel free to give it a shot on the playgrounds. https://fal.ai/models…

Congrats Burkay - the model is very impressive. One area I’d like to see improved in a flux v2 is knowledge of artist styles. Flux cannot respond to requests asking for paintings in the style of David Hockney, Norman Rockwell, Edgar Degas, — it seems to have no fine art training at all. I’d bet that fine art training would further improve the compositional skills of the model, plus it would open up a range of uses th…

>Flux cannot respond to requests asking for paintings in the style of David Hockney, Norman Rockwell

Does it respond to any names? I noticed SD3 removed all names to prevent recreating famous people but as a side effect lost the very powerful ability to infer styles from artist names too.

Re: Flux: Open-source text-to-image model with 12B parameters

#170
post #79

Earlier quoted context omitted.

On my list of AI concerns, whether or not Vanity Fair has it’s copyright infringed does not appear.

First they came for fashion magazines, and I said nothing.

All journalism is just duplicating the works and performance of others without their permission for profit anyway.
Post reply on HN