Live data from Hacker News

VASA-1: Lifelike audio-driven talking faces generated in real time

microsoft.com

31–40 of 166 posts

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#31

“We have no plans to release an online demo, API, product, additional implementation details, or any related offerings until we are certain that the technology will be used responsibly and in accordance with proper regulations.”

Translation: "We're attempting to preserve our moat, and this is the correct PR blurb. We'll release an API once we're far enough ahead and extracted enough money."

Like somebody on Ars noted "anybody notice it's an election year?" You don't need to release an API, all online videos are now suspicious authenticity. Somebody make a video of Trump or Biden's eyes following the mouse cursor around. Real videos turned into fake videos.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#33
post #32

If you see talking heads with static/simple/blurred backgrounds from now on, assume it is fake. In the near future they will accompany realistic backgrounds and even less detectable fakes, we will have to assume all vids could be faked.

I still find the faces themselves to be really obviously wrong. The sound is just off, close enough to tell who is being imitated but not particularly good.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#34
post #33
post #32

If you see talking heads with static/simple/blurred backgrounds from now on, assume it is fake. In the near future they will accompany realistic backgrounds and even less detectable fakes, we will have to assume all vids could be faked.

I still find the faces themselves to be really obviously wrong. The sound is just off, close enough to tell who is being imitated but not particularly good.

Especially the hair "physics" and sometimes the teeth shift around a bit.

But that's nitpicking. It's good enough to fool someone not watching too closely. And the fact that the result is this good with a single photo is truly astonishing, we used to have to train models on thousands of photos for days only to end up with a worse result!

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#35
post #32

If you see talking heads with static/simple/blurred backgrounds from now on, assume it is fake. In the near future they will accompany realistic backgrounds and even less detectable fakes, we will have to assume all vids could be faked.

I wonder how video evidence in court is going to be affected by this. Both from a defense and prosecution perspective.

Technically videos could've been faked before but it would require a ton of effort and skill that no average person would have.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#36
What this is starting to reveal is that there's a clear need for some kind of chain of custody system that guarantees the authenticity of what we see. Nikon/Canon tried doing this in the past, but improper storage of private keys lead to vulnerabilities. As far as I'm aware it's never extended to video either.

With modern secure hardware keys it may yet be possible. The difficulty is that any kind of photo/video manipulation would break the signature (and there are practical reasons to want to be able to edit videos obviously).

In the ideal world, any mutation to the original source content could be traceable to the original source content. But that's not an easy problem to solve.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#38
I'm curious what is the reason for deepfake research and what the practical application is.

Can someone explain the commercial need to take someones likeness and generate video content?

If I was an a-list celebrity, I would give permission for coke to make a commercial with my likeness, provided I am allowed final approval of the finished ad?

Do I have an avatar that attends my zoom work calls?

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#39

I'm curious what is the reason for deepfake research and what the practical application is. Can someone explain the commercial need to take someones likeness and generate video content? If I was an a-list celebrity, I would give permission for coke to make a commercial with my likeness, provided I am allowed final approval of the finished ad? Do I have an avatar that attends my zoom work calls?

State disinformation and propaganda campaigns.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#40

What this is starting to reveal is that there's a clear need for some kind of chain of custody system that guarantees the authenticity of what we see. Nikon/Canon tried doing this in the past, but improper storage of private keys lead to vulnerabilities. As far as I'm aware it's never extended to video either. With modern secure hardware keys it may yet be possible. The difficulty is that any kind of photo/video mani…

No, we are merely returning to the pre-photography state of things where a mere printed image is not sufficient evidence for anything.
Post reply on HN