Live data from Hacker News

SnapFusion: Text-to-Image Diffusion Model on Mobile Devices Within Two Seconds

snap-research.github.io

51–54 of 54 posts

Re: SnapFusion: Text-to-Image Diffusion Model on Mobile Devices Within Two Seconds

#51
post #34

Earlier quoted context omitted.

Asking for evidence isn’t a “far fetched theory.” A video isn’t evidence. I’m looking forward to more info proving the claims are accurate.

Videos are absolutely evidence. You’re looking for someone to replicate their findings, which is fine, but your basis for believing it to be possibly fraudulent is far-fetched.

Not OP, but I guess "nullius in verba" has gone out of fashion then?

Re: SnapFusion: Text-to-Image Diffusion Model on Mobile Devices Within Two Seconds

#52

Earlier quoted context omitted.

The paper describes the implementation including a detailed breakdown of the optimisation algorithm itself. It's also plausible an iPhone 14 Pro could do it given its memory b/w, ops/s and that it can fit the SD model in RAM.

Hopefully someone will replicate it, and then we’ll know if they were telling the truth.

So you just assume these smart people are all lying ? And not just like hiding details but straight up saying it's on-device but faking it ? If so, I don't understand why that is your reaction.

Re: SnapFusion: Text-to-Image Diffusion Model on Mobile Devices Within Two Seconds

#53

From the paper: > In this work, we present the first text-to-image diffusion model that generates an image on mobile devices in less than 2 seconds. To achieve this, we mainly focus on improving the slow inference speed of the UNet and reducing the number of necessary denoising steps. As a layman, it's impressive and surprising that there's so much room for optimization here, given the number of hands on folks in the…

I'm sad that Carmack decided to (as I understand it) focus away from LLMs because they're "already getting enough eyes" -- it feels like he wants to make a novel paradigm shift kind of contribution, but his magic power has always seemed to me to be a capability of grasping a huge amount of technical depth in detail, and seeing past the easy local optima to the real essence of the computation being done, and finding ways to measure and squeeze everything out of that.

Carmack could possibly get us realtime networked stable diffusion text to video and video to video at high resolution, maybe even on phones. It will probably happen anyway, but it might take 5+ extra years, and there'll probably be a ton of stupid things we never fix.

Re: SnapFusion: Text-to-Image Diffusion Model on Mobile Devices Within Two Seconds

#54
post #46

I would rather have more quality than speed. The output of this model reminds me of Midjourney 3.

There are lots of models/approaches to look at if you want improved speed. It seems bizarre to not also want the speed to increase.

I would absolutely want more speed if the quality was comparable to the competitor's offerings. But it looks like it's a speed up at the expense of the image quality.
Post reply on HN