Live data from Hacker News

SnapFusion: Text-to-Image Diffusion Model on Mobile Devices Within Two Seconds

snap-research.github.io

11–20 of 54 posts

Re: SnapFusion: Text-to-Image Diffusion Model on Mobile Devices Within Two Seconds

#11
> Text-to-image diffusion models can create stunning images from natural language descriptions that rival the work of professional artists and photographers.

I’m all for the continued advance of diffusion models.

If this paper offered evidence of quantitative and qualities measurement techniques for determining human preference for art or photos based on a prompt, I’d get it the phrasing.

But having the first sentence essentially spurn professional creatives seems to unnecessarily fan the flames.

AI image generation does bother some creatives, and there are real reasons for this given the many models that have been trained using long practiced work.

Keep up with the science, but don’t forget the tact!

Re: SnapFusion: Text-to-Image Diffusion Model on Mobile Devices Within Two Seconds

#14
post #2

This is insane! But it makes me wonder if we've reached a local maximum in AI where the current methods are great at generating still images but they're pretty much uncontrollable. Like if you ask the AI to generate a dog, is it really possible to prompt every single detail so it creates exactly what you have in mind, or is it more like a trust situation where you just accept whatever the AI generates for you?

It's a human problem really. If you asked an artist to draw a dog you'd have to "trust" them? To control every detail you'd have to tell them every detail, either upfront (whereby the artist night struggle to achieve) or as a series of edits. In the latter case both artist and AI would struggle to keep the look consistent the more edits you make.

At some point you might as well just do it yourself cause it’s easier. I’ve gotten to this point more than once with ChatGPT. ChatGPT will get you like 70% of the way initially and if you are lucky you’ll hit 90% with a lot of time invested in “prompt engineering”.

The thing is, none of these are mind readers. And text is a very poor way to define tight specs. The best way for software is to code it yourself. Code is the spec. Same with drawing. The spec is the drawing itself. Only the human can control that.

…or something…

Re: SnapFusion: Text-to-Image Diffusion Model on Mobile Devices Within Two Seconds

#15
Why focus on mobile?

Near real-time rendering while entering a prompt on desktop would be even more amazing!

Imagine e.g. UI sliders for adding weight to multi-prompts like "night time" (ideally sliding from [day time]:1, [night time]:0 to both zero, to all night with 0,1).

Re: SnapFusion: Text-to-Image Diffusion Model on Mobile Devices Within Two Seconds

#17

Why focus on mobile? Near real-time rendering while entering a prompt on desktop would be even more amazing! Imagine e.g. UI sliders for adding weight to multi-prompts like "night time" (ideally sliding from [day time]:1, [night time]:0 to both zero, to all night with 0,1).

It has a total latency of 200ms on a A100 40G.

Re: SnapFusion: Text-to-Image Diffusion Model on Mobile Devices Within Two Seconds

#18
Sub 2 second generations on cell phones, nice! Better FID and CLIP scores than Stable Diffusion v1.5 with 50 steps, great!

So are they gonna release the code, or do they only open-source ad-SDKs[1]?

[1] https://github.com/orgs/Snapchat/repositories

Re: SnapFusion: Text-to-Image Diffusion Model on Mobile Devices Within Two Seconds

#19

Why focus on mobile? Near real-time rendering while entering a prompt on desktop would be even more amazing! Imagine e.g. UI sliders for adding weight to multi-prompts like "night time" (ideally sliding from [day time]:1, [night time]:0 to both zero, to all night with 0,1).

> Why focus on mobile?

Snap Inc.

Re: SnapFusion: Text-to-Image Diffusion Model on Mobile Devices Within Two Seconds

#20

Why focus on mobile? Near real-time rendering while entering a prompt on desktop would be even more amazing! Imagine e.g. UI sliders for adding weight to multi-prompts like "night time" (ideally sliding from [day time]:1, [night time]:0 to both zero, to all night with 0,1).

> Why focus on mobile? Snap Inc.

Oh you're so right. The face filter industrial complex is going to eat this up. And after them, plastic surgeons... God help us all.
Post reply on HN