Live data from Hacker News

VASA-1: Lifelike audio-driven talking faces generated in real time

microsoft.com

131–140 of 166 posts

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#131

Why is this research being done? Is this some kind of arms race? The only purpose of this technology I can think of is getting spies to abuse others. Am I going to have to do AuthN and AuthZ on every phone call and zoom now?

We all know why this is really happening. Clippy 2.0.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#132
post #54

Earlier quoted context omitted.

On the other hand, if deepfaking becomes common enough that everyone stops trusting everything they read / see on the internet, it would be a net good against the spread of disinformation compared to today.

I don't see that as an outcome. We have already seen a grand erosion of trust in institutions. Moving to an even lower trust society does not sound like it would have positive consequences for discourse, public policy, or society at large.

The benefit is that you can only trust in person interaction with social and governmental institutions so people will have to leave their damn house again and go talk to each other face to face. Too many of our current problems are caused by people only interacting with each other and the world through third parties who are performing a MITM operation for their own benefit.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#133

Oh god don't watch their teeth! Proper creepy. Still, apart from the teeth this looks extremely convincing!

The teeth resizing dynamically is incredibly distracting, or more positively, a nice way to identify fakes. For now.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#134

Earlier quoted context omitted.

On the other hand, if deepfaking becomes common enough that everyone stops trusting everything they read / see on the internet, it would be a net good against the spread of disinformation compared to today.

I don't see the extinction of trust through the introduction of garbage falsehoods to be a net good. Believing that everything you eat is poisoned is no way to live. Believing that everything you see is a lie is also no way to live.

Before photography this was just the normal state of the world. Think a little, back then any story or picture you saw was made by a person and you only had their reputation to go by. Think some more and you realize that’s never changed even with pictures and video. Easy AI generated pictures and video just remove the illusion of trust.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#135

What this is starting to reveal is that there's a clear need for some kind of chain of custody system that guarantees the authenticity of what we see. Nikon/Canon tried doing this in the past, but improper storage of private keys lead to vulnerabilities. As far as I'm aware it's never extended to video either. With modern secure hardware keys it may yet be possible. The difficulty is that any kind of photo/video mani…

I think it's practically impossible for such a system to be globally trustworthy due to the practical inevitability of "but improper storage of private keys lead to vulnerabilities" scenarios. People will expect or require that chain of custody only if all or at least the vast majority of the content they want would have that chain of custody. Photo/video content will have that chain of custody only if all or almost…

Multisig by the user and camera manufacturer can help to some extent.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#136

Why is this research being done? Is this some kind of arms race? The only purpose of this technology I can think of is getting spies to abuse others. Am I going to have to do AuthN and AuthZ on every phone call and zoom now?

Newscasters and other talking heads will be out of business. Just pipe the script into some AI and get video.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#137

It looks all warpy and stretchy. That's not how skin and face muscles work. Looks fake to me.

I find the hairs to be the least realistic, they look elastic, which is unsurprising: highly detailed things like hairs are hard to simulate with good fidelity.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#138
post #69

And it's only going to get faster, better, easier, cheaper.[a] Meanwhile, yesterday my credit card company asked me if I wanted to use voice authentication for verifying my identity "more securely" on the phone. Surely the company spent many millions of dollars to enable this new security-theater feature. It begs the question: Is every single executive and manager at my credit card company completely unaware that rig…

Yes, they are and they also know it isn't foolproof so that isn't the only information being compared against. Some services compare the calling number is compared against live activity on the PSTN (ie, subscriber's phone is not in an active call, but their number is being presented as as the caller ID is one such metric). Many of these deep fake generators with public access have watermarks in the audio. The audio stream comparison continues, it needs to speak like you, word and phrase choices. There are other fingerprints of generated audio that you can't hear, but are still obvious at the moment. With security, it always cat and mouse with fraudsters on one hand and the effort/frustration with customers on the other.

Asking customers questions that they don't remember and that fraudsters have in front of them isn't working and the time it takes for agents to authenticate is very expensive.

While there is no doubt that companies will screw up with security, you are making wild accusations without reference to any evidence.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#139
post #131

Why is this research being done? Is this some kind of arms race? The only purpose of this technology I can think of is getting spies to abuse others. Am I going to have to do AuthN and AuthZ on every phone call and zoom now?

We all know why this is really happening. Clippy 2.0.

[deleted]

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#140
post #69

And it's only going to get faster, better, easier, cheaper.[a] Meanwhile, yesterday my credit card company asked me if I wanted to use voice authentication for verifying my identity "more securely" on the phone. Surely the company spent many millions of dollars to enable this new security-theater feature. It begs the question: Is every single executive and manager at my credit card company completely unaware that rig…

Any time you add a "new" security gate to your product, it should be in addition to and not instead of the existing gates. Biometrics should not replace username/password, they should be in addition to. Security Questions like "What was your first pet's name" should not be able to get you in the backdoor. SMS verification alone should not allow you to reset your password. Same with this voice authentication stuff. It…

If you make your product sufficiently inconvenient, then you'll have the unassailable security posture of having no users.
Post reply on HN