Live data from Hacker News

FLUX.2 [Klein]: Towards Interactive Visual Intelligence

bfl.ai

31–40 of 59 posts

Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence

#31
post #27

Earlier quoted context omitted.

They are underselling Z-Image Turbo somewhat. It's arguably the best overall model for local image generation for several reasons including prompt adherence, overall output quality and realism, and freedom from censorship, even though it's also one of the smallest at 6B parameters. ZIT is not far short of revolutionary. It is kind of surreal to contemplate how much high-quality imagery can be extracted from a model t…

Hold on now. Z-Image Turbo has gotten a lot of hype but it's worse at all of those things other than perhaps looking like it was shot on a cell phone camera than Qwen Image and Flux 2 (the full sized version). Once you get away from photographic portraits of people it quickly shows just how little it can do. It is, however, small and quick.

Not in my experience. Flux 2 is much larger and heavily censored, and Qwen-Image is just plain not as good. You can fool me into thinking that Z-Image Turbo output isn't AI, while that's rarely the case with Qwen.

Look at the images I posted elsewhere in this section. They are crappy excuses for pogo sticks, but they absolutely do NOT look like they came from a cell phone.

Also see vunderba's page at https://genai-showdown.specr.net/ . Even when Z-Image Turbo fails a test, it still looks great most of the time.

Edit re: your other comment -- don't make the mistake of confusing censorship with lack of training data. Z-Image will try to render whatever you ask for, but at the end of the day it's a very small model that will fail once you start asking for things it simply wasn't trained on. They didn't train it with much NSFW material, so it has some rather... unorthodox anatomical ideas.

Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence

#33

I haven’t gotten around to adding Klein to my GenAI Showdown site yet, but if it’s anything like Z-Image Turbo, it should perform extremely well. For reference, Z-Image Turbo scored 4 out of 15 points on GenAI Showdown. I’m aware that doesn’t sound like much, but given that one of the largest models, Flux.2 (32b), only managed to outscore ZiT (a 6b model) by a single point and is significantly heavier-weight, that’s…

Can you fix the information bubble on mobile please? When pressing one, it vanishes instantly...

Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence

#34
post #19

It cannot create an image of a pogo stick. I was trying to get it to create an image of a tiger jumping on a pogo stick, which is way beyond its capabilities, but it cannot create an image of a pogo stick in isolation.

You are right, just tried even with reference images it can't do it for me. Maybe with some good prompting.

Because in theory I would say that knowledge is something that does not have to be baked in the model but could be added using reference images if the model is capable enough to reason about them.

Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence

#35

I haven’t gotten around to adding Klein to my GenAI Showdown site yet, but if it’s anything like Z-Image Turbo, it should perform extremely well. For reference, Z-Image Turbo scored 4 out of 15 points on GenAI Showdown. I’m aware that doesn’t sound like much, but given that one of the largest models, Flux.2 (32b), only managed to outscore ZiT (a 6b model) by a single point and is significantly heavier-weight, that’s…

Can you fix the information bubble on mobile please? When pressing one, it vanishes instantly...

Hey Bombthecat, sorry about that! I can't repro this issue on any of the devices I have (Android Pixel 7, an iPad, etc).

If you get a chance, could you list your mobile device specs? That way I can at least try it on Browserstack and see if I can figure out a fix.

Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence

#36

If we think of GenAI models as a compression implementation. Generally, text compresses extremely well. Images and video do not. Yet state-of-the-art text-to-image and text-to-video models are often much smaller (in parameter count) than large language models like Llama-3. Maybe vision models are small because we’re not actually compressing very much of the visual world. The training data covers a narrow, human-biase…

I find it likely that we are still missing a few major efficiency tricks with LLMs. But I would also not underestimate the amount of implicit knowledge and skill an LLM is expected to carry on a meta level.

Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence

#37

Earlier quoted context omitted.

Can you fix the information bubble on mobile please? When pressing one, it vanishes instantly...

Hey Bombthecat, sorry about that! I can't repro this issue on any of the devices I have (Android Pixel 7, an iPad, etc). If you get a chance, could you list your mobile device specs? That way I can at least try it on Browserstack and see if I can figure out a fix.

Samsung, brave browser

Update: Huh, now it's working

Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence

#38
post #19

It cannot create an image of a pogo stick. I was trying to get it to create an image of a tiger jumping on a pogo stick, which is way beyond its capabilities, but it cannot create an image of a pogo stick in isolation.

When given an image of an empty wine glass, it can't fill it to the brim with wine. The pogo stick drawers and wine glass fillers can enjoy their job security for months to come!

Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence

#40
post #38
post #19

It cannot create an image of a pogo stick. I was trying to get it to create an image of a tiger jumping on a pogo stick, which is way beyond its capabilities, but it cannot create an image of a pogo stick in isolation.

When given an image of an empty wine glass, it can't fill it to the brim with wine. The pogo stick drawers and wine glass fillers can enjoy their job security for months to come!

You can still taste wine in the metaverse with the mouth adapter and can get a buzz by gently electrifying your neuralink (time travel required)
Post reply on HN