I am amazed, though not entirely surprised, that these models keep getting smaller while the quality and effectiveness increases. z image turbo is wild, I'm looking forward to trying this one out. An older thread on this has a lot of comments: https://news.ycombinator.com/item?id=46046916
Quality is increasing, but these small models have very little knowledge compared to their big brothers (Qwen Image/Full size Flux 2). As in characters, artists, specific items, etc.
FLUX.2 [Klein]: Towards Interactive Visual Intelligence
41–50 of 59 posts
Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence
#42Flux2 Klein isn’t some generation leap or anything. It’s good, but let’s be honest, this is an ad. What will be really interesting to me is the release of Z-image, if that goes the way it’s looking, it’ll be natural language SDXL 2.0, which seems to be what people really want. Releasing the Turbo/Distilled/Finetune months ago was a genius move really. It hurt Flux and Qwen releases on a possible future implication al…
I think that information still did not get through to most users.
"Notably, the resulting distilled model not only matches the original multi-step teacher but even surpasses it in terms of photorealism and visual impact."
"It achieves 8-step inference that is not only indistinguishable from the 100-step teacher but frequently surpasses it in perceived quality and aesthetic appeal"
Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence
#43Flux2 Klein isn’t some generation leap or anything. It’s good, but let’s be honest, this is an ad. What will be really interesting to me is the release of Z-image, if that goes the way it’s looking, it’ll be natural language SDXL 2.0, which seems to be what people really want. Releasing the Turbo/Distilled/Finetune months ago was a genius move really. It hurt Flux and Qwen releases on a possible future implication al…
The team behind Z-Image Turbo has told us multiple times in their paper that the output quality of the Turbo model is superior to the larger base model. I think that information still did not get through to most users. "Notably, the resulting distilled model not only matches the original multi-step teacher but even surpasses it in terms of photorealism and visual impact." "It achieves 8-step inference that is not onl…
Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence
#44Earlier quoted context omitted.
The team behind Z-Image Turbo has told us multiple times in their paper that the output quality of the Turbo model is superior to the larger base model. I think that information still did not get through to most users. "Notably, the resulting distilled model not only matches the original multi-step teacher but even surpasses it in terms of photorealism and visual impact." "It achieves 8-step inference that is not onl…
It's important for finetuning, Lora training and as a refiner...
However, I wonder what has been the source of the delay with its release and if there were problems with that approach.
Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence
#45I am amazed, though not entirely surprised, that these models keep getting smaller while the quality and effectiveness increases. z image turbo is wild, I'm looking forward to trying this one out. An older thread on this has a lot of comments: https://news.ycombinator.com/item?id=46046916
Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence
#46I haven’t gotten around to adding Klein to my GenAI Showdown site yet, but if it’s anything like Z-Image Turbo, it should perform extremely well. For reference, Z-Image Turbo scored 4 out of 15 points on GenAI Showdown. I’m aware that doesn’t sound like much, but given that one of the largest models, Flux.2 (32b), only managed to outscore ZiT (a 6b model) by a single point and is significantly heavier-weight, that’s…
Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence
#47If we think of GenAI models as a compression implementation. Generally, text compresses extremely well. Images and video do not. Yet state-of-the-art text-to-image and text-to-video models are often much smaller (in parameter count) than large language models like Llama-3. Maybe vision models are small because we’re not actually compressing very much of the visual world. The training data covers a narrow, human-biase…
> Generally, text compresses extremely well. Images and video do not. Is that actually true? I'm not sure it's fair to compare lossless compression ratios of text (abstract, noiseless) to images and video that innately have random sampling noise. If you look at humanly indistinguishable compression, I'd expect that you'd see far better compression ratios for lossy image and video compression than lossless text.
Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence
#48I am amazed, though not entirely surprised, that these models keep getting smaller while the quality and effectiveness increases. z image turbo is wild, I'm looking forward to trying this one out. An older thread on this has a lot of comments: https://news.ycombinator.com/item?id=46046916
Quality is increasing, but these small models have very little knowledge compared to their big brothers (Qwen Image/Full size Flux 2). As in characters, artists, specific items, etc.
And small models are also much easier to fine tune than large ones.
Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence
#49Re: FLUX.2 [Klein]: Towards Interactive Visual Intelligence
#50Earlier quoted context omitted.
Can you fix the information bubble on mobile please? When pressing one, it vanishes instantly...
Hey Bombthecat, sorry about that! I can't repro this issue on any of the devices I have (Android Pixel 7, an iPad, etc). If you get a chance, could you list your mobile device specs? That way I can at least try it on Browserstack and see if I can figure out a fix.