Live data from Hacker News

VASA-1: Lifelike audio-driven talking faces generated in real time

microsoft.com

21–30 of 166 posts

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#22
The paper mentions it uses Diffusion Transformers. The open source implementation that comes up in Google is Facebook Research's PyTorch implementation which is a non-commercial license. https://github.com/facebookresearch/DiT

Is there something equivalent but MIT or Apache?

I feel like diffusion transformers are key now.

I wonder if OpenAI implemented their SORA stuff from scratch or if they built on the Facebook Research diffusion transformers library. That would be interesting if they violated the non-commercial part.

Hm. Found one: https://github.com/milmor/diffusion-transformer-keras

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#23

Oh god don't watch their teeth! Proper creepy. Still, apart from the teeth this looks extremely convincing!

yeah, teeth, tongue movement and lack of tongue shape and the "stretching" of the skin around the cheeks in the images pushed the videos right into the uncanny valley for me.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#24
post #11

This is absolutely crazy. And it'll only get better from here. Imagine "VASA-9" or whatever. I thought deepfakes were still quite a bit away but after this I will have to be way more careful online. It's not far from behind something that can show up in your "YouTube shorts" feed and trick you if you didn't already know it was AI.

This is good but nowhere as good as EMO https://humanaigc.github.io/emote-portrait-alive/ ( https://news.ycombinator.com/item?id=39533326 ) This one has too much movement and looks eerie/robotic/uncanny valley. While EMO looks just perfect.

Hard disagree -- I think you might be misremembering how EMO looks in practice -- I'm sure we'll learn VASA-1 "telltales" but to my eyes there are far fewer than EMO - zero of the EMO videos were 'perfect' for me, and many show little glitches or missing sync. VASA-1 still blinks a bit more than I think is natural, but it looks much more fluid.

Both are, BTW, AMAZING!! Pretty crazy.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#25

“We have no plans to release an online demo, API, product, additional implementation details, or any related offerings until we are certain that the technology will be used responsibly and in accordance with proper regulations.”

> until we are certain that the technology will be used responsibly ... That's basically "never" then, so we'll see how long they hold out. Scammers are already using the existing voice/image/video generation apparently fairly successfully. :(

Having a delay, where people can see what's coming down the pipe, does have value. In a year there may/will be a open source model.

But knowing that this is possible is important to know.

I'm fairly clued in, and am constantly surprised at how fast things are changing.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#26

Earlier quoted context omitted.

> until we are certain that the technology will be used responsibly ... That's basically "never" then, so we'll see how long they hold out. Scammers are already using the existing voice/image/video generation apparently fairly successfully. :(

Having a delay, where people can see what's coming down the pipe, does have value. In a year there may/will be a open source model. But knowing that this is possible is important to know. I'm fairly clued in, and am constantly surprised at how fast things are changing.

> But knowing that this is possible ...

Who knowing this is possible?

The general elderly person isn't going to know any time soon. The SV IT people probably will.

It's not an even distribution of knowledge. ;/

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#29
post #12

My first thought was "oh no the interview fakes", but then I realized - what if they just kept using the face? Would I care?

Yeah, even if they just use LLMs to do all the work, or are a LLM themselves, as long as they can do the work I guess.

Weird implications for various regulations though.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#30

Earlier quoted context omitted.

This is good but nowhere as good as EMO https://humanaigc.github.io/emote-portrait-alive/ ( https://news.ycombinator.com/item?id=39533326 ) This one has too much movement and looks eerie/robotic/uncanny valley. While EMO looks just perfect.

Hard disagree -- I think you might be misremembering how EMO looks in practice -- I'm sure we'll learn VASA-1 "telltales" but to my eyes there are far fewer than EMO - zero of the EMO videos were 'perfect' for me, and many show little glitches or missing sync. VASA-1 still blinks a bit more than I think is natural, but it looks much more fluid. Both are, BTW, AMAZING!! Pretty crazy.

In VASA there is way to much body movement instead of just being he head as if camera is moving in the strong winds. EMO is a lot more human like. In the very first video on the EMO page I still cannot see it as a generated video, its that real. The lip movement, the expressions are in almost in perfect sync with the voice. That is absolutely not the case with VASA
Post reply on HN