Live data from Hacker News

VASA-1: Lifelike audio-driven talking faces generated in real time

microsoft.com

101–110 of 166 posts

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#101
post #74

Anyone have any good ideas for how we're going to do politics now? Today a big ML model can do this and it's somewhat regulate-able, tomorrow people can do this on their contact-lens supercomputers and anyone can generate a video of anything. Is going back to personally knowing your local representative the only way? How will we vote for national candidates if nobody knows what they think or say?

> Today a big ML model can do this Not that big: https://github.com/Zejun-Yang/AniPortrait https://huggingface.co/ZJYang/AniPortrait/tree/main

Didn't see that one pretty cool, not as good as Emo or Vasa but pretty good

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#102

Why is this research being done? Is this some kind of arms race? The only purpose of this technology I can think of is getting spies to abuse others. Am I going to have to do AuthN and AuthZ on every phone call and zoom now?

Advertising. Now you and your friends star in the streaming commercials and digital billboards near you! (whether you want to or not)

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#103

Why is this research being done? Is this some kind of arms race? The only purpose of this technology I can think of is getting spies to abuse others. Am I going to have to do AuthN and AuthZ on every phone call and zoom now?

On the other hand, if deepfaking becomes common enough that everyone stops trusting everything they read / see on the internet, it would be a net good against the spread of disinformation compared to today.

everyone stops trusting everything

Why would you expect this to happen? Lots of people are gullible, if it were otherwise a lot of well-known politicians would be out of a job or would never have been elected to begin with.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#104

What this is starting to reveal is that there's a clear need for some kind of chain of custody system that guarantees the authenticity of what we see. Nikon/Canon tried doing this in the past, but improper storage of private keys lead to vulnerabilities. As far as I'm aware it's never extended to video either. With modern secure hardware keys it may yet be possible. The difficulty is that any kind of photo/video mani…

No, we are merely returning to the pre-photography state of things where a mere printed image is not sufficient evidence for anything.

merely

You say this as if it were not a big deal, but losing a century's worth of authentication infrastructure/practises is a Bad Thing which will have large negative externalities.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#105

Anyone have any good ideas for how we're going to do politics now? Today a big ML model can do this and it's somewhat regulate-able, tomorrow people can do this on their contact-lens supercomputers and anyone can generate a video of anything. Is going back to personally knowing your local representative the only way? How will we vote for national candidates if nobody knows what they think or say?

People already believe any quote you slap on a JPEG.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#106

Anyone have any good ideas for how we're going to do politics now? Today a big ML model can do this and it's somewhat regulate-able, tomorrow people can do this on their contact-lens supercomputers and anyone can generate a video of anything. Is going back to personally knowing your local representative the only way? How will we vote for national candidates if nobody knows what they think or say?

> Anyone have any good ideas for how we're going to do politics now?

If a business is showing a demo of this you can be assured that the Government already has this tech and has for a period of time.

> How will we vote for national candidates if nobody knows what they think or say?

You don't know what they think or say now - hopefully this disabuses people of this notion.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#107
post #69

And it's only going to get faster, better, easier, cheaper.[a] Meanwhile, yesterday my credit card company asked me if I wanted to use voice authentication for verifying my identity "more securely" on the phone. Surely the company spent many millions of dollars to enable this new security-theater feature. It begs the question: Is every single executive and manager at my credit card company completely unaware that rig…

Any time you add a "new" security gate to your product, it should be in addition to and not instead of the existing gates. Biometrics should not replace username/password, they should be in addition to. Security Questions like "What was your first pet's name" should not be able to get you in the backdoor. SMS verification alone should not allow you to reset your password. Same with this voice authentication stuff. It should be another layer, not a replacement of your actual credentials.

If you treat it as OR instead of AND, then your security is only as good as the worst link in the chain.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#108

Earlier quoted context omitted.

On the other hand, if deepfaking becomes common enough that everyone stops trusting everything they read / see on the internet, it would be a net good against the spread of disinformation compared to today.

everyone stops trusting everything Why would you expect this to happen? Lots of people are gullible, if it were otherwise a lot of well-known politicians would be out of a job or would never have been elected to begin with.

If it's even commoner than "common enough" then anyone could at least try to help their gullible friends and family by sending them a deepfake video of them doing/saying something they've never said. A lot of people will suddenly wise up when a problem affects them directly.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#109
> To show off the model, Microsoft created a VASA-1 research page featuring many sample videos of the tool in action

With AI stuff, I have learned to be very skeptical until and unless a relatively publicly accessible demo with user specified inputs is available.

It is way too easy for humans to cherry pick the nice outputs, or to take advantage of biases in the training data to generate nice outputs, and is not at all reflective of how it holds up in the real world.

Part of the reason why ChatGPT, Stable Diffusion, Dall-E had such an impact is the people could try and see for themselves without being told how awesome it was by the people making it.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#110
post #88

Anyone have any good ideas for how we're going to do politics now? Today a big ML model can do this and it's somewhat regulate-able, tomorrow people can do this on their contact-lens supercomputers and anyone can generate a video of anything. Is going back to personally knowing your local representative the only way? How will we vote for national candidates if nobody knows what they think or say?

We already rely on chains of trust going back to the original source, and will still. I find these alarmist posts a bit mystifying – before photography, anyone could fake a quote of anyone, and human civilisation got quite far. We had a bit over a hundred years where phographic-quality images were possible and very hard to fake (which did and still does vary with technology), but clearly now we’re past that. We’ll ma…

The issue is better phrased as “how will we survive the transition while some folk still believe the video they are seeing is irrefutable proof the event happened?”
Post reply on HN