Live data from Hacker News

VASA-1: Lifelike audio-driven talking faces generated in real time

microsoft.com

141–150 of 166 posts

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#141
post #54

Earlier quoted context omitted.

I don't see that as an outcome. We have already seen a grand erosion of trust in institutions. Moving to an even lower trust society does not sound like it would have positive consequences for discourse, public policy, or society at large.

The benefit is that you can only trust in person interaction with social and governmental institutions so people will have to leave their damn house again and go talk to each other face to face. Too many of our current problems are caused by people only interacting with each other and the world through third parties who are performing a MITM operation for their own benefit.

This assumes that it's a two-way door.

Over the past century and a half, we've moved into vast, anonymous spaces, where I'm as likely to know and get along with my neighbour as I am to win the lottery.

And this is important. No, it's not just a matter of putting on an effort to learn who my neighbour is -- my neighbour is literally someone whose life experiences are wildly different, whose social outcomes will be wildly different, whose beliefs and values are wildly different, and, for all I know, goes to conferences about how to eliminate me and my kind.

(This last part is not speculation; I'm trans; see: CPAC)

And these are my reasons. My neighbour is probably equivalently terrified of me, or what I represent, or the media I consume, or the conferences that I go to.

Generalizing, you can't take a bunch of random people whose only bond is that they share meatspace-proximity, draw a circle around them, and declare them a community; those communities are _gone_, and you can no more bring them back than you can revive a corpse. (This would also probably not be a good idea, even if it were possible: they were also incredibly uncomfortable places for anyone who didn't fit in, and we have generations of fiction about people risking everything to leave for those big anonymous cities we created in step 1.)

So, here we are, dependent on technology to stay in touch with far-flung friends and lovers and family, all of us, scattered like spiderwebs across the globe, and now into the strands drips a poison.

Daniel Dennett was right. Counterfeit people are an enormous danger to civilization. Research like this should stop immediately.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#142

Why is this research being done? Is this some kind of arms race? The only purpose of this technology I can think of is getting spies to abuse others. Am I going to have to do AuthN and AuthZ on every phone call and zoom now?

> Why is this research being done?

I think it's mostly "because it can be done". These types of impressive demos have become relatively low hanging fruit in terms of how modern machine learning can be applied.

One could imagine commercial applications (VR, virtual "try before you buy", etc), but things like this can also be a flex by the AI labs, or a PhD student wanting to write a paper.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#144

Earlier quoted context omitted.

I think it's practically impossible for such a system to be globally trustworthy due to the practical inevitability of "but improper storage of private keys lead to vulnerabilities" scenarios. People will expect or require that chain of custody only if all or at least the vast majority of the content they want would have that chain of custody. Photo/video content will have that chain of custody only if all or almost…

Multisig by the user and camera manufacturer can help to some extent.

Multisig requires user cooperation, many users will not care to cooperate, and chain of custody verification really starts working only if you can get (force) ~100% of legitimate users globally to adopt the system.

Also, for the potential creators of political fakes, such a multisig won't change things - getting a manufacturer's key may take some effort, but getting (and 'burning') keys of a dozen random people is relatively trivial in many ways - e.g. buying off of poor people, stealing from compromised random machines, or simply issuing fake identities for state-backed actors.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#145

Anyone have any good ideas for how we're going to do politics now? Today a big ML model can do this and it's somewhat regulate-able, tomorrow people can do this on their contact-lens supercomputers and anyone can generate a video of anything. Is going back to personally knowing your local representative the only way? How will we vote for national candidates if nobody knows what they think or say?

> Anyone have any good ideas for how we're going to do politics now? If a business is showing a demo of this you can be assured that the Government already has this tech and has for a period of time. > How will we vote for national candidates if nobody knows what they think or say? You don't know what they think or say now - hopefully this disabuses people of this notion.

> If a business is showing a demo of this you can be assured that the Government already has this tech and has for a period of time.

That may have been true once upon a time, but it no longer is. And even in the areas it was true it was mostly for niche areas like cryptanalysis.

Governments simply cannot attract or keep the level of talent required to have been far ahead of industry on LLMs and similar tech, especially not with the huge difference in salaries and working conditions.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#146

Why is this research being done? Is this some kind of arms race? The only purpose of this technology I can think of is getting spies to abuse others. Am I going to have to do AuthN and AuthZ on every phone call and zoom now?

I was thinking about this the other day. An implantable Yubikey type device that integrates with whatever device you’re using to validate your identity for phone calls or video conferences.

Subdermal X.509 maybe with some sort of neurolink adapter so you can confirm the request for identity. Though, first versions might be just a small button you need to press during the handshake.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#147
post #69

And it's only going to get faster, better, easier, cheaper.[a] Meanwhile, yesterday my credit card company asked me if I wanted to use voice authentication for verifying my identity "more securely" on the phone. Surely the company spent many millions of dollars to enable this new security-theater feature. It begs the question: Is every single executive and manager at my credit card company completely unaware that rig…

The point is lowering liability. By choosing to not use voice authentication (or whatever), it becomes easier to argue that fraud is your fault. Or if you did use it, the company 'is doing everything they can' and 'exceeding industry standards' so it isn't their fault, either. It also just makes them seem more secure to the uninitiated (the security-theater bit, yes).

Maybe one day someone will successfully argue that adding easily defeated checks lowers security, by adding friction for no reason or instilling false confidence in users at both ends.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#148

Anyone have any good ideas for how we're going to do politics now? Today a big ML model can do this and it's somewhat regulate-able, tomorrow people can do this on their contact-lens supercomputers and anyone can generate a video of anything. Is going back to personally knowing your local representative the only way? How will we vote for national candidates if nobody knows what they think or say?

> Anyone have any good ideas for how we're going to do politics now? If a business is showing a demo of this you can be assured that the Government already has this tech and has for a period of time. > How will we vote for national candidates if nobody knows what they think or say? You don't know what they think or say now - hopefully this disabuses people of this notion.

[deleted]

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#149
post #48

Earlier quoted context omitted.

There goes the dashcam industry…

You're being downvoted but I think the comment raises a good question. what will happen when someone gets accused of doctoring their dashcam footage? Or any footage used for evidence.

I wasn’t really kidding with my comment. I just recently used camera footage as part of an accident claim and the assessor immediately said “that wasn’t your fault, we take responsibility on behalf of the driver”.

In a few years time when (if) faking realistic footage becomes trivial, I suspect this kind of video will have a much, much higher level of scrutiny or only be accepted from certain sources such as government owned traffic cameras.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#150

What this is starting to reveal is that there's a clear need for some kind of chain of custody system that guarantees the authenticity of what we see. Nikon/Canon tried doing this in the past, but improper storage of private keys lead to vulnerabilities. As far as I'm aware it's never extended to video either. With modern secure hardware keys it may yet be possible. The difficulty is that any kind of photo/video mani…

None of that works, it's simply theatre.

I can just take a (crypto-signed) photo of another photo.

Post reply on HN