Live data from Hacker News

VASA-1: Lifelike audio-driven talking faces generated in real time

microsoft.com

121–130 of 166 posts

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#121
post #63

Earlier quoted context omitted.

True, an image, audio clip, or video is not enough evidence to establish truth. We still need a way to establish truth. It's important for security cameras, for politics, and for public figures. Here are some things we could start looking into. * Cameras that sign their output. Yes, this camera caught this video, and it hasn't been modified. This is a must for recordings being used in court evidence IMO. Otherwise fr…

Blockchains can be used for cryptography time-stamping. I’ve always had a suspicion that governments and large companies would prefer a world without hard cryptographic proofs. After wikileaks they noticed DKIM can cause them major blowback. Somehow general public isn’t aware all the emails were proven authentic with DKIM signatures and even in fairly educated circles people believe the “emails were fake” but it’s no…

Quite the opposite, governments and large companies even explicitly run services for digital timestamping of documents - if I wanted to potentially assert some facts in court, I'd definitely prefer having that e-document with a timestamp notarized from my local government service instead of Bitcoin, because while the cryptography is the same, it would be much simpler from the practical legal perspective, requiring less time and effort and cost to get the court to accept that.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#122
post #12

My first thought was "oh no the interview fakes", but then I realized - what if they just kept using the face? Would I care?

It would be interesting that a remote candidate could easily identify as whatever ethnicity, age or even gender they consider most beneficial for hiring to avoid discrimination or fit certain diversity incentives.

Tech like this has the potential to bring us back to the days of "on the Internet, nobody know's you're a dog" https://en.wikipedia.org/wiki/On_the_Internet,_nobody_knows_...

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#123

Earlier quoted context omitted.

> Is every single executive and manager at my credit card company completely unaware that right now anyone can clone anyone else's voice by obtaining a short sample audio clip taken from any social network? Your mistake is assuming the company cares. The "company" is a hundred different disjointed departments that only care about not getting caught Equifax-style (or filing for bankruptcy if caught). If the marketing…

Also, this feature is probably just some midd level execs plan for a bonus, not a rigorously reviewed and planned. It's also probably in the pipeline for a decade so if they don't push it out, suddenly they get no bonus for cancelling a project. Corporations are ultimately no better than governments and likely worse depending on what their regulatory environment looks like.

There’s a really important thing here for anyone trying to do sales to big companies.

Find an exec that needs a project to advance their career. Make your software that project.

Suck in as many other execs into the project so their careers become coupled to getting your software rolled out.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#124

Earlier quoted context omitted.

Also, this feature is probably just some midd level execs plan for a bonus, not a rigorously reviewed and planned. It's also probably in the pipeline for a decade so if they don't push it out, suddenly they get no bonus for cancelling a project. Corporations are ultimately no better than governments and likely worse depending on what their regulatory environment looks like.

There’s a really important thing here for anyone trying to do sales to big companies. Find an exec that needs a project to advance their career. Make your software that project. Suck in as many other execs into the project so their careers become coupled to getting your software rolled out.

That's clever!

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#125
post #69

And it's only going to get faster, better, easier, cheaper.[a] Meanwhile, yesterday my credit card company asked me if I wanted to use voice authentication for verifying my identity "more securely" on the phone. Surely the company spent many millions of dollars to enable this new security-theater feature. It begs the question: Is every single executive and manager at my credit card company completely unaware that rig…

That scene from Sneakers would be so different nowadays. [1]

"My voice is my passport. Verify me." [2]

1. https://youtu.be/WdcIqFOc2UE?si=Df3DtSakatp9eD0L

2. https://youtu.be/-zVgWpVXb64?si=yT2GZpb7E2yZoEYl

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#126
post #33
post #32

If you see talking heads with static/simple/blurred backgrounds from now on, assume it is fake. In the near future they will accompany realistic backgrounds and even less detectable fakes, we will have to assume all vids could be faked.

I still find the faces themselves to be really obviously wrong. The sound is just off, close enough to tell who is being imitated but not particularly good.

It's interesting to me that some of the long-standing things are still there. For example, lots of people with an earring in only one ear, unlikely asymmetry in the shape or size of their ears, etc.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#128

What this is starting to reveal is that there's a clear need for some kind of chain of custody system that guarantees the authenticity of what we see. Nikon/Canon tried doing this in the past, but improper storage of private keys lead to vulnerabilities. As far as I'm aware it's never extended to video either. With modern secure hardware keys it may yet be possible. The difficulty is that any kind of photo/video mani…

I expect this type of system to be implemented in my lifetime. It will allow whistleblowers and investigative sources to be discredited or tracked down and persecuted.

Unfortunately that seems inevitable.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#129

Earlier quoted context omitted.

No, we are merely returning to the pre-photography state of things where a mere printed image is not sufficient evidence for anything.

merely You say this as if it were not a big deal, but losing a century's worth of authentication infrastructure/practises is a Bad Thing which will have large negative externalities.

It isn't really though. It has been technical possible to convincingly doctor photos for some time already, gradually getting easier, cheaper, and faster with time for decades, and even now the currently available tech has limitations and the full change is not going to happen overnight.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#130
post #69

And it's only going to get faster, better, easier, cheaper.[a] Meanwhile, yesterday my credit card company asked me if I wanted to use voice authentication for verifying my identity "more securely" on the phone. Surely the company spent many millions of dollars to enable this new security-theater feature. It begs the question: Is every single executive and manager at my credit card company completely unaware that rig…

[deleted]
Post reply on HN