Live data from Hacker News

VASA-1: Lifelike audio-driven talking faces generated in real time

microsoft.com

111–120 of 166 posts

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#111
post #88

Anyone have any good ideas for how we're going to do politics now? Today a big ML model can do this and it's somewhat regulate-able, tomorrow people can do this on their contact-lens supercomputers and anyone can generate a video of anything. Is going back to personally knowing your local representative the only way? How will we vote for national candidates if nobody knows what they think or say?

We already rely on chains of trust going back to the original source, and will still. I find these alarmist posts a bit mystifying – before photography, anyone could fake a quote of anyone, and human civilisation got quite far. We had a bit over a hundred years where phographic-quality images were possible and very hard to fake (which did and still does vary with technology), but clearly now we’re past that. We’ll ma…

In the before times we didn't have social media and its algorithms and reach. Does it matter that the chains of trust debunk a viral lie 24 hours after it had spread? Not that there's a lot of trust in the chains of trust to begin with. And if you still have trust, then you're not the target of the viral lie. And if you still have trust, then how long can you hold on that trust when the lies keep coming 24/7 one after another without end. As one movie critic once put it: You might not have noticed it, but your brain did. Very malleable this brain of ours.

The civilization might be fine, sure. Now, democracy, on the other hand...

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#114

Why is this research being done? Is this some kind of arms race? The only purpose of this technology I can think of is getting spies to abuse others. Am I going to have to do AuthN and AuthZ on every phone call and zoom now?

On the other hand, if deepfaking becomes common enough that everyone stops trusting everything they read / see on the internet, it would be a net good against the spread of disinformation compared to today.

That's the whole issue though, spread of disinformation eroded trust, furthering this into obliteration of all trust is not a good outcome.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#115
post #69

And it's only going to get faster, better, easier, cheaper.[a] Meanwhile, yesterday my credit card company asked me if I wanted to use voice authentication for verifying my identity "more securely" on the phone. Surely the company spent many millions of dollars to enable this new security-theater feature. It begs the question: Is every single executive and manager at my credit card company completely unaware that rig…

> Is every single executive and manager at my credit card company completely unaware that right now anyone can clone anyone else's voice by obtaining a short sample audio clip taken from any social network? Your mistake is assuming the company cares. The "company" is a hundred different disjointed departments that only care about not getting caught Equifax-style (or filing for bankruptcy if caught). If the marketing…

[dead]

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#116
post #64

Why is this research being done? Is this some kind of arms race? The only purpose of this technology I can think of is getting spies to abuse others. Am I going to have to do AuthN and AuthZ on every phone call and zoom now?

Because the text for this is only a slight variation of the tech for a broad range of legitimate applications? Because even this precise tech has legitimate use cases? > The only purpose of this technology I can think of is getting spies to abuse others. Can you really not think of any other use cases?

Why don't you list some legitimate and useful values of this work? Especially at the price we and this company are paying.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#117
post #69

And it's only going to get faster, better, easier, cheaper.[a] Meanwhile, yesterday my credit card company asked me if I wanted to use voice authentication for verifying my identity "more securely" on the phone. Surely the company spent many millions of dollars to enable this new security-theater feature. It begs the question: Is every single executive and manager at my credit card company completely unaware that rig…

I mean, what do you want them to do? If we think their security officers are freaking out and holding meetings right now about what to do, or if they're asleep at the wheel, we'd be seeing the same thing from the outside, no?

No, because multiple companies are pushing this atm. If it was only company I would agree, but with multiple, you'd have at least one that would back out of it again.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#119

What this is starting to reveal is that there's a clear need for some kind of chain of custody system that guarantees the authenticity of what we see. Nikon/Canon tried doing this in the past, but improper storage of private keys lead to vulnerabilities. As far as I'm aware it's never extended to video either. With modern secure hardware keys it may yet be possible. The difficulty is that any kind of photo/video mani…

I think it's practically impossible for such a system to be globally trustworthy due to the practical inevitability of "but improper storage of private keys lead to vulnerabilities" scenarios.

People will expect or require that chain of custody only if all or at least the vast majority of the content they want would have that chain of custody.

Photo/video content will have that chain of custody only if all or almost all of devices recording that content will support it - including all the cheapest mass-produced devices in reasonably widespread use anywhere in the world.

And that chain of custody provides the benefit only if literally 100% of these manufacturers have their private keys secure 100% of the time, which is simply not happening; at least one such key will leak, if not unintentionally then intentionally for some intelligence agency who wants to fake content.

And what do you do once you see a leak of the private keys used for signing the certificates for the private keys securely embedded in (for example) all of 2029 Huawei smartphones, which could be like 200 million phones? The users won't replace their phones just because of that, and you'll have all these users making content - so everyone will have to choose to either auto-block and discard everything from all those 200 million users, or permit content with a potentially fake chain of custody; and I'm totally certain that most people will prefer the latter.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#120
post #32

If you see talking heads with static/simple/blurred backgrounds from now on, assume it is fake. In the near future they will accompany realistic backgrounds and even less detectable fakes, we will have to assume all vids could be faked.

I wonder how video evidence in court is going to be affected by this. Both from a defense and prosecution perspective. Technically videos could've been faked before but it would require a ton of effort and skill that no average person would have.

Just as before, a major part of photo or video evidence in court is not the actual video itself, but a person testifying "on that day I saw this horrible event, where these things happened, and here's attached evidence that I filmed which illustrates some details of what I saw." - which would be a valid consideration even without the photo/video, but the added details do obviously help.

Courts already wouldn't generally approve random footage without clear provenance.

Post reply on HN