Live data from Hacker News

VASA-1: Lifelike audio-driven talking faces generated in real time

microsoft.com

51–60 of 166 posts

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#51

Why is this research being done? Is this some kind of arms race? The only purpose of this technology I can think of is getting spies to abuse others. Am I going to have to do AuthN and AuthZ on every phone call and zoom now?

I get the feeling it's "someone's going to do this, so it might as well be us."

It's fascinating how research can take on a life of its own and will be pushed, by someone, to its own conclusion. Even for immensely destructive technologies (e.g., atomic weapons, viruses), the impact of a technology is its own attractor (could you say that's risk-seeking behavior?)

> Am I going to have to do AuthN and AuthZ on every phone call and zoom now?

"Alexa, I need an alibi for yesterday at noon."

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#52
post #32

If you see talking heads with static/simple/blurred backgrounds from now on, assume it is fake. In the near future they will accompany realistic backgrounds and even less detectable fakes, we will have to assume all vids could be faked.

I wonder how video evidence in court is going to be affected by this. Both from a defense and prosecution perspective. Technically videos could've been faked before but it would require a ton of effort and skill that no average person would have.

There will be a new cottage industry of AI detectives that serve as expert witnesses and they will attest to the originality of media to the court

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#53
post #48

Earlier quoted context omitted.

No, we are merely returning to the pre-photography state of things where a mere printed image is not sufficient evidence for anything.

There goes the dashcam industry…

You're being downvoted but I think the comment raises a good question. what will happen when someone gets accused of doctoring their dashcam footage? Or any footage used for evidence.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#54

Why is this research being done? Is this some kind of arms race? The only purpose of this technology I can think of is getting spies to abuse others. Am I going to have to do AuthN and AuthZ on every phone call and zoom now?

On the other hand, if deepfaking becomes common enough that everyone stops trusting everything they read / see on the internet, it would be a net good against the spread of disinformation compared to today.

I don't see that as an outcome. We have already seen a grand erosion of trust in institutions. Moving to an even lower trust society does not sound like it would have positive consequences for discourse, public policy, or society at large.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#57

What this is starting to reveal is that there's a clear need for some kind of chain of custody system that guarantees the authenticity of what we see. Nikon/Canon tried doing this in the past, but improper storage of private keys lead to vulnerabilities. As far as I'm aware it's never extended to video either. With modern secure hardware keys it may yet be possible. The difficulty is that any kind of photo/video mani…

I expect this type of system to be implemented in my lifetime. It will allow whistleblowers and investigative sources to be discredited or tracked down and persecuted.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#58

I'm curious what is the reason for deepfake research and what the practical application is. Can someone explain the commercial need to take someones likeness and generate video content? If I was an a-list celebrity, I would give permission for coke to make a commercial with my likeness, provided I am allowed final approval of the finished ad? Do I have an avatar that attends my zoom work calls?

Video games, entertainment, and avatars seems like the big ones.

If that is really the reason then this is insane and everyone involved should put their keyboards down and stop what they are doing.

This would be as if we invented and sold nuclear weapons to dig out quarry mines faster. The inconvenience it saves us quickly disappears into the overwhelming shadow of the enormous harm now enabled.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#59

I'm curious what is the reason for deepfake research and what the practical application is. Can someone explain the commercial need to take someones likeness and generate video content? If I was an a-list celebrity, I would give permission for coke to make a commercial with my likeness, provided I am allowed final approval of the finished ad? Do I have an avatar that attends my zoom work calls?

Imagine being the CEO and you just grab your salary and options, go home, sit in the hot tub while one of the interns carefully prompts GPT and VASE how you are giving a speech online about strategic directions. /s
Post reply on HN