Earlier quoted context omitted.
Asking for evidence isn’t a “far fetched theory.” A video isn’t evidence. I’m looking forward to more info proving the claims are accurate.
Videos are absolutely evidence. You’re looking for someone to replicate their findings, which is fine, but your basis for believing it to be possibly fraudulent is far-fetched.
SnapFusion: Text-to-Image Diffusion Model on Mobile Devices Within Two Seconds
51–54 of 54 posts
Re: SnapFusion: Text-to-Image Diffusion Model on Mobile Devices Within Two Seconds
#52Earlier quoted context omitted.
The paper describes the implementation including a detailed breakdown of the optimisation algorithm itself. It's also plausible an iPhone 14 Pro could do it given its memory b/w, ops/s and that it can fit the SD model in RAM.
Hopefully someone will replicate it, and then we’ll know if they were telling the truth.
Re: SnapFusion: Text-to-Image Diffusion Model on Mobile Devices Within Two Seconds
#53From the paper: > In this work, we present the first text-to-image diffusion model that generates an image on mobile devices in less than 2 seconds. To achieve this, we mainly focus on improving the slow inference speed of the UNet and reducing the number of necessary denoising steps. As a layman, it's impressive and surprising that there's so much room for optimization here, given the number of hands on folks in the…
Carmack could possibly get us realtime networked stable diffusion text to video and video to video at high resolution, maybe even on phones. It will probably happen anyway, but it might take 5+ extra years, and there'll probably be a ton of stupid things we never fix.
Re: SnapFusion: Text-to-Image Diffusion Model on Mobile Devices Within Two Seconds
#54I would rather have more quality than speed. The output of this model reminds me of Midjourney 3.
There are lots of models/approaches to look at if you want improved speed. It seems bizarre to not also want the speed to increase.