Still, apart from the teeth this looks extremely convincing!
VASA-1: Lifelike audio-driven talking faces generated in real time
21–30 of 166 posts
Re: VASA-1: Lifelike audio-driven talking faces generated in real time
#22Is there something equivalent but MIT or Apache?
I feel like diffusion transformers are key now.
I wonder if OpenAI implemented their SORA stuff from scratch or if they built on the Facebook Research diffusion transformers library. That would be interesting if they violated the non-commercial part.
Hm. Found one: https://github.com/milmor/diffusion-transformer-keras
Re: VASA-1: Lifelike audio-driven talking faces generated in real time
#23Oh god don't watch their teeth! Proper creepy. Still, apart from the teeth this looks extremely convincing!
Re: VASA-1: Lifelike audio-driven talking faces generated in real time
#24This is absolutely crazy. And it'll only get better from here. Imagine "VASA-9" or whatever. I thought deepfakes were still quite a bit away but after this I will have to be way more careful online. It's not far from behind something that can show up in your "YouTube shorts" feed and trick you if you didn't already know it was AI.
This is good but nowhere as good as EMO https://humanaigc.github.io/emote-portrait-alive/ ( https://news.ycombinator.com/item?id=39533326 ) This one has too much movement and looks eerie/robotic/uncanny valley. While EMO looks just perfect.
Both are, BTW, AMAZING!! Pretty crazy.
Re: VASA-1: Lifelike audio-driven talking faces generated in real time
#25“We have no plans to release an online demo, API, product, additional implementation details, or any related offerings until we are certain that the technology will be used responsibly and in accordance with proper regulations.”
> until we are certain that the technology will be used responsibly ... That's basically "never" then, so we'll see how long they hold out. Scammers are already using the existing voice/image/video generation apparently fairly successfully. :(
But knowing that this is possible is important to know.
I'm fairly clued in, and am constantly surprised at how fast things are changing.
Re: VASA-1: Lifelike audio-driven talking faces generated in real time
#26Earlier quoted context omitted.
> until we are certain that the technology will be used responsibly ... That's basically "never" then, so we'll see how long they hold out. Scammers are already using the existing voice/image/video generation apparently fairly successfully. :(
Having a delay, where people can see what's coming down the pipe, does have value. In a year there may/will be a open source model. But knowing that this is possible is important to know. I'm fairly clued in, and am constantly surprised at how fast things are changing.
Who knowing this is possible?
The general elderly person isn't going to know any time soon. The SV IT people probably will.
It's not an even distribution of knowledge. ;/
Re: VASA-1: Lifelike audio-driven talking faces generated in real time
#27real jurassic park "too preoccupied with whether they could" vibes
Re: VASA-1: Lifelike audio-driven talking faces generated in real time
#28i get why this is interesting but why is it desirable? real jurassic park "too preoccupied with whether they could" vibes
Re: VASA-1: Lifelike audio-driven talking faces generated in real time
#29My first thought was "oh no the interview fakes", but then I realized - what if they just kept using the face? Would I care?
Weird implications for various regulations though.
Re: VASA-1: Lifelike audio-driven talking faces generated in real time
#30Earlier quoted context omitted.
This is good but nowhere as good as EMO https://humanaigc.github.io/emote-portrait-alive/ ( https://news.ycombinator.com/item?id=39533326 ) This one has too much movement and looks eerie/robotic/uncanny valley. While EMO looks just perfect.
Hard disagree -- I think you might be misremembering how EMO looks in practice -- I'm sure we'll learn VASA-1 "telltales" but to my eyes there are far fewer than EMO - zero of the EMO videos were 'perfect' for me, and many show little glitches or missing sync. VASA-1 still blinks a bit more than I think is natural, but it looks much more fluid. Both are, BTW, AMAZING!! Pretty crazy.