Live data from Hacker News

FLUX1.1 [pro] – New SotA text-to-image model from Black Forest Labs

replicate.com

91–100 of 159 posts

Re: FLUX1.1 [pro] – New SotA text-to-image model from Black Forest Labs

#91

I'm running Asahi Linux on a 32GB M1 Pro. Any chance of being able to run text-to-image models locally? I've had some success with LLMs, but only the smaller models. No idea where to start with images, everything seems geared towards msft+nvda.

Try https://github.com/leejet/stable-diffusion.cpp

Re: FLUX1.1 [pro] – New SotA text-to-image model from Black Forest Labs

#92

It doesn’t get piano keyboards right, but it’s the first image generator I’ve tried that sometimes get “someone playing accordion” mostly right. When I ask for a man playing accordion, it’s usually a somewhat flawed piano accordion, but If I ask for a woman playing accordion, it’s usually a button accordion. I’ve also seen a few that are half-button, half-piano monstrosities. Also, if I ask for “someone playing accor…

Periodic data is always hard for generative image systems - particularly if that "cycle" window is relatively large (as would be the case for octaves of a piano).

Yeah, it's my informal test to see if a new model has made any progress on that.

Re: FLUX1.1 [pro] – New SotA text-to-image model from Black Forest Labs

#93
post #48

Pretty smart model. Here's one I made: https://replicate.com/p/6ez0x8xqvsrga0cjadg8m7bah0

Yet, it doesn't seem to know how a Tektronix 4010 actually looks like... ;) I had similar issues trying to paint a "I cast non-magic missile" meme with a fantasy wizard using a missile launcher. No model out there (I've tried SD, SDXL, FLUX.1dev and now this FLUX1.1pro) knows how a missile launcher looks like (neither as a generic term, nor any specific systems) and even has no clue how it's held, so they all draw re…

Isn't it because the shoulder launched weapon is usually called rocket launcher, rpg or bazooka? Never heard it referred as misille launcher.

Re: FLUX1.1 [pro] – New SotA text-to-image model from Black Forest Labs

#94
post #48

Pretty smart model. Here's one I made: https://replicate.com/p/6ez0x8xqvsrga0cjadg8m7bah0

It's quite good at following a detailed paragraph long description of an scene, which is a double edged sword. A lot of the fun for me with early text to image models was underspecifying an image and then enjoying how the model "invents" it. "Steampunk spaceship", "communist bear", "glass city".

flux is amazing, but I find it requires a very literal description, which pushes the "creative work" back to the text itself. Which can certainly be a good thing, just a bit less gratifying to non visual types like myself. :)

I wonder, only somewhat jokingly, if one could make text generators which "imagine" detailed fantastical scenes, suitable for feeding to a text to image model.

Re: FLUX1.1 [pro] – New SotA text-to-image model from Black Forest Labs

#95

I'm running Asahi Linux on a 32GB M1 Pro. Any chance of being able to run text-to-image models locally? I've had some success with LLMs, but only the smaller models. No idea where to start with images, everything seems geared towards msft+nvda.

DiffusionBee will let you do this quite easily. edit: nevermind, it's a macos app

Is DiffusionBee still in development? I had stopped using it because it seemed like the dev interest had stalled.

Re: FLUX1.1 [pro] – New SotA text-to-image model from Black Forest Labs

#96
post #94
post #48

Pretty smart model. Here's one I made: https://replicate.com/p/6ez0x8xqvsrga0cjadg8m7bah0

It's quite good at following a detailed paragraph long description of an scene, which is a double edged sword. A lot of the fun for me with early text to image models was underspecifying an image and then enjoying how the model "invents" it. "Steampunk spaceship", "communist bear", "glass city". flux is amazing, but I find it requires a very literal description, which pushes the "creative work" back to the text itsel…

That's what Fooocus is - it allows you to specify a "text expander" LLM that sits in between the input prompt and the diffusion model.

https://github.com/lllyasviel/Fooocus

Re: FLUX1.1 [pro] – New SotA text-to-image model from Black Forest Labs

#99
post #48

Pretty smart model. Here's one I made: https://replicate.com/p/6ez0x8xqvsrga0cjadg8m7bah0

One thing that makes FLUX so special is the prompt understanding. I now gave FLUX 1.1 a prompt "Closeup of a doll house built to resemble a famous room in the TV show Friends" and it gave me one with the sign "Central Perk". I never prompted for the text "Central Perk". A Redditor also discovered that it has an associative understanding of emotions. For example "Rose of passion" and it may draw a flower that is burning, because passion is fiery.

This is miles ahead of most other image generation models available today.

Re: FLUX1.1 [pro] – New SotA text-to-image model from Black Forest Labs

#100
post #97

Earlier quoted context omitted.

Have you run Ideogram offline?

Have you run Flux Pro offline?

No, only a dozen Flux Dev models different distillations, quantizations, and fine-tunes with LORAs.

But you keep pretending that close source AI is a sustainable comparison.

Post reply on HN