Live data from Hacker News

FLUX.2 [Klein]: Towards Interactive Visual Intelligence

bfl.ai

41–50 of 59 posts

Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence

#41
post #26
post #2

I am amazed, though not entirely surprised, that these models keep getting smaller while the quality and effectiveness increases. z image turbo is wild, I'm looking forward to trying this one out. An older thread on this has a lot of comments: https://news.ycombinator.com/item?id=46046916

Quality is increasing, but these small models have very little knowledge compared to their big brothers (Qwen Image/Full size Flux 2). As in characters, artists, specific items, etc.

I smell the bias-variance tradeoff. By underfitting more, they get closer to the degenerate case of a model that only knows one perfect photo.

Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence

#42

Flux2 Klein isn’t some generation leap or anything. It’s good, but let’s be honest, this is an ad. What will be really interesting to me is the release of Z-image, if that goes the way it’s looking, it’ll be natural language SDXL 2.0, which seems to be what people really want. Releasing the Turbo/Distilled/Finetune months ago was a genius move really. It hurt Flux and Qwen releases on a possible future implication al…

The team behind Z-Image Turbo has told us multiple times in their paper that the output quality of the Turbo model is superior to the larger base model.

I think that information still did not get through to most users.

"Notably, the resulting distilled model not only matches the original multi-step teacher but even surpasses it in terms of photorealism and visual impact."

"It achieves 8-step inference that is not only indistinguishable from the 100-step teacher but frequently surpasses it in perceived quality and aesthetic appeal"

https://arxiv.org/abs/2511.22699

Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence

#43

Flux2 Klein isn’t some generation leap or anything. It’s good, but let’s be honest, this is an ad. What will be really interesting to me is the release of Z-image, if that goes the way it’s looking, it’ll be natural language SDXL 2.0, which seems to be what people really want. Releasing the Turbo/Distilled/Finetune months ago was a genius move really. It hurt Flux and Qwen releases on a possible future implication al…

The team behind Z-Image Turbo has told us multiple times in their paper that the output quality of the Turbo model is superior to the larger base model. I think that information still did not get through to most users. "Notably, the resulting distilled model not only matches the original multi-step teacher but even surpasses it in terms of photorealism and visual impact." "It achieves 8-step inference that is not onl…

It's important for finetuning, Lora training and as a refiner...

Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence

#44

Earlier quoted context omitted.

The team behind Z-Image Turbo has told us multiple times in their paper that the output quality of the Turbo model is superior to the larger base model. I think that information still did not get through to most users. "Notably, the resulting distilled model not only matches the original multi-step teacher but even surpasses it in terms of photorealism and visual impact." "It achieves 8-step inference that is not onl…

It's important for finetuning, Lora training and as a refiner...

I also heard so, that it would mainly be useful for training and applying the resulting Lora to the distilled Turbo model.

However, I wonder what has been the source of the delay with its release and if there were problems with that approach.

Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence

#45
post #2

I am amazed, though not entirely surprised, that these models keep getting smaller while the quality and effectiveness increases. z image turbo is wild, I'm looking forward to trying this one out. An older thread on this has a lot of comments: https://news.ycombinator.com/item?id=46046916

Is there a theoritical minimum for params for a given output? I saw news about GPT 3.5, then Deepseek training models at a fraction of that cost, then laptops running a model that beats 3.5. When does it stop?

Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence

#46

I haven’t gotten around to adding Klein to my GenAI Showdown site yet, but if it’s anything like Z-Image Turbo, it should perform extremely well. For reference, Z-Image Turbo scored 4 out of 15 points on GenAI Showdown. I’m aware that doesn’t sound like much, but given that one of the largest models, Flux.2 (32b), only managed to outscore ZiT (a 6b model) by a single point and is significantly heavier-weight, that’s…

I think it shows problems with your tests tbh. The bigger models are way more capable than you make them out to be. They are also better in training and understanding of CGI render outputs as reference like normal maps or id-masks. Your testing suite is the perfect example that structured data implies false confidence. Pure t2i is not a good benchmark anymore.

Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence

#47

If we think of GenAI models as a compression implementation. Generally, text compresses extremely well. Images and video do not. Yet state-of-the-art text-to-image and text-to-video models are often much smaller (in parameter count) than large language models like Llama-3. Maybe vision models are small because we’re not actually compressing very much of the visual world. The training data covers a narrow, human-biase…

> Generally, text compresses extremely well. Images and video do not. Is that actually true? I'm not sure it's fair to compare lossless compression ratios of text (abstract, noiseless) to images and video that innately have random sampling noise. If you look at humanly indistinguishable compression, I'd expect that you'd see far better compression ratios for lossy image and video compression than lossless text.

The comparison makes sense in what I am charitably assuming is the case the GP is referring to: we know how to build a tight embedding space from a text corpus, and get out outputs from it tolerably similar to the inputs for the purposes they're put to. That is lossy compression, just not in the sense anyone talking about conventional lossless text compression algorithms would use the words. I'm not sure we can say the same of image embeddings.

Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence

#48
post #26
post #2

I am amazed, though not entirely surprised, that these models keep getting smaller while the quality and effectiveness increases. z image turbo is wild, I'm looking forward to trying this one out. An older thread on this has a lot of comments: https://news.ycombinator.com/item?id=46046916

Quality is increasing, but these small models have very little knowledge compared to their big brothers (Qwen Image/Full size Flux 2). As in characters, artists, specific items, etc.

That's what LoRAs are for.

And small models are also much easier to fine tune than large ones.

Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence

#50

Earlier quoted context omitted.

Can you fix the information bubble on mobile please? When pressing one, it vanishes instantly...

Hey Bombthecat, sorry about that! I can't repro this issue on any of the devices I have (Android Pixel 7, an iPad, etc). If you get a chance, could you list your mobile device specs? That way I can at least try it on Browserstack and see if I can figure out a fix.

Yeah works fine for me on a Pixel 9.
Post reply on HN