Live data from Hacker News

VASA-1: Lifelike audio-driven talking faces generated in real time

microsoft.com

61–70 of 166 posts

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#61
Anyone have any good ideas for how we're going to do politics now?

Today a big ML model can do this and it's somewhat regulate-able, tomorrow people can do this on their contact-lens supercomputers and anyone can generate a video of anything.

Is going back to personally knowing your local representative the only way? How will we vote for national candidates if nobody knows what they think or say?

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#63

What this is starting to reveal is that there's a clear need for some kind of chain of custody system that guarantees the authenticity of what we see. Nikon/Canon tried doing this in the past, but improper storage of private keys lead to vulnerabilities. As far as I'm aware it's never extended to video either. With modern secure hardware keys it may yet be possible. The difficulty is that any kind of photo/video mani…

No, we are merely returning to the pre-photography state of things where a mere printed image is not sufficient evidence for anything.

True, an image, audio clip, or video is not enough evidence to establish truth.

We still need a way to establish truth. It's important for security cameras, for politics, and for public figures. Here are some things we could start looking into.

* Cameras that sign their output. Yes, this camera caught this video, and it hasn't been modified. This is a must for recordings being used in court evidence IMO. Otherwise framing a crime is as easy as a few deep fakes and planting some DNA or fingerprints at the scene of the crime.

* People digitally signing pictures/audio/videos of them. Even if they digitally modified the data it shows that they consent to having their image associated with that message. It reduces the strength of the attack vector of deep fake videos for reputation sabotage.

* Malicious content source detection and flagging. Think email spam filter type tagging of fake content. Community notes on X would be another good example.

* Digital manipulation detection. I'm less than hopeful this will be the way in the long term, but could be used to disprove some fraud.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#64

Why is this research being done? Is this some kind of arms race? The only purpose of this technology I can think of is getting spies to abuse others. Am I going to have to do AuthN and AuthZ on every phone call and zoom now?

Because the text for this is only a slight variation of the tech for a broad range of legitimate applications?

Because even this precise tech has legitimate use cases?

> The only purpose of this technology I can think of is getting spies to abuse others.

Can you really not think of any other use cases?

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#66

Anyone have any good ideas for how we're going to do politics now? Today a big ML model can do this and it's somewhat regulate-able, tomorrow people can do this on their contact-lens supercomputers and anyone can generate a video of anything. Is going back to personally knowing your local representative the only way? How will we vote for national candidates if nobody knows what they think or say?

People in my circles have been saying this for a few years now, and we've yet to see it happen.

I've got my popcorn ready.

But you can rest easy. Everyone just votes for the candidate their party picked, anyway.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#67
post #54

Earlier quoted context omitted.

On the other hand, if deepfaking becomes common enough that everyone stops trusting everything they read / see on the internet, it would be a net good against the spread of disinformation compared to today.

I don't see that as an outcome. We have already seen a grand erosion of trust in institutions. Moving to an even lower trust society does not sound like it would have positive consequences for discourse, public policy, or society at large.

Ironically low effort deep fakes might increase trust in organizations that have had the budget to fake stuff since their inception. The losers are 'citizen journalist' broadcasting on Youtube etc.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#68

I'm curious what is the reason for deepfake research and what the practical application is. Can someone explain the commercial need to take someones likeness and generate video content? If I was an a-list celebrity, I would give permission for coke to make a commercial with my likeness, provided I am allowed final approval of the finished ad? Do I have an avatar that attends my zoom work calls?

In this case, replacing humans in service jobs. From the paper:

"Such technology holds the promise of enriching digital communication, increasing accessibility for those with communicative impairments, transforming education methods with interactive AI tutoring, and providing therapeutic support and social interaction in healthcare."

A convincing simulacrum of empathy could plausibly be the most profitable product since oil.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#69
And it's only going to get faster, better, easier, cheaper.[a]

Meanwhile, yesterday my credit card company asked me if I wanted to use voice authentication for verifying my identity "more securely" on the phone. Surely the company spent many millions of dollars to enable this new security-theater feature.

It begs the question: Is every single executive and manager at my credit card company completely unaware that right now anyone can clone anyone else's voice by obtaining a short sample audio clip taken from any social network? If anyone is aware, why is the company acting like this?

Corporate America is so far behind the times it's not even funny.

---

[a] With apologies to Daft Punk.

Post reply on HN