Live data from Hacker News

VASA-1: Lifelike audio-driven talking faces generated in real time

microsoft.com

91–100 of 166 posts

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#93
post #88

Anyone have any good ideas for how we're going to do politics now? Today a big ML model can do this and it's somewhat regulate-able, tomorrow people can do this on their contact-lens supercomputers and anyone can generate a video of anything. Is going back to personally knowing your local representative the only way? How will we vote for national candidates if nobody knows what they think or say?

We already rely on chains of trust going back to the original source, and will still. I find these alarmist posts a bit mystifying – before photography, anyone could fake a quote of anyone, and human civilisation got quite far. We had a bit over a hundred years where phographic-quality images were possible and very hard to fake (which did and still does vary with technology), but clearly now we’re past that. We’ll ma…

Yeah I mean tabloids have been fooling people with doctored photos for decades.

Potentially we'll need slightly tighter regulations on formal press (so that people that care for accurate information have a place they can get it) and definitely we'll want to steer the culture back towards holding them accountable for misinformation, but credulous people have always had easy access to bad information.

I'm much more worried at the potential abuse cases that involve ordinary people that aren't public figures, and have much less ability to defend themselves. Heck, even celebrities are a more vulnerable targets than politicians.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#94

What this is starting to reveal is that there's a clear need for some kind of chain of custody system that guarantees the authenticity of what we see. Nikon/Canon tried doing this in the past, but improper storage of private keys lead to vulnerabilities. As far as I'm aware it's never extended to video either. With modern secure hardware keys it may yet be possible. The difficulty is that any kind of photo/video mani…

No, we are merely returning to the pre-photography state of things where a mere printed image is not sufficient evidence for anything.

Pre-photography it at least took effort, practice, and time, to draw something convincing. Any skill with that much of a barrier to entry kind of automatically reduces the ability to be anonymous. And we didn't have the ability to instantaneously distribute images world-wide.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#95
post #88

Anyone have any good ideas for how we're going to do politics now? Today a big ML model can do this and it's somewhat regulate-able, tomorrow people can do this on their contact-lens supercomputers and anyone can generate a video of anything. Is going back to personally knowing your local representative the only way? How will we vote for national candidates if nobody knows what they think or say?

We already rely on chains of trust going back to the original source, and will still. I find these alarmist posts a bit mystifying – before photography, anyone could fake a quote of anyone, and human civilisation got quite far. We had a bit over a hundred years where phographic-quality images were possible and very hard to fake (which did and still does vary with technology), but clearly now we’re past that. We’ll ma…

Presidential elections are frequently pretty close. Taking the electoral college into account (not the popular vote, which doesn't matter) Donald Trump won the 2016 election by a grand total of ~80,000 votes in three states[0].

Knowing that retractions rarely get viral exposure, it's not difficult to imagine that a few sufficiently-viral videos could swing enough votes to impact a presidential election. Especially when considering that the average person is not up to speed on the current state of the tech, and so has not been prompted to build up the mindset that's required to fend off this new threat.

[0] https://www.washingtonpost.com/news/the-fix/wp/2016/12/01/do...

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#96
post #88

Earlier quoted context omitted.

We already rely on chains of trust going back to the original source, and will still. I find these alarmist posts a bit mystifying – before photography, anyone could fake a quote of anyone, and human civilisation got quite far. We had a bit over a hundred years where phographic-quality images were possible and very hard to fake (which did and still does vary with technology), but clearly now we’re past that. We’ll ma…

Presidential elections are frequently pretty close. Taking the electoral college into account (not the popular vote, which doesn't matter) Donald Trump won the 2016 election by a grand total of ~80,000 votes in three states[0]. Knowing that retractions rarely get viral exposure, it's not difficult to imagine that a few sufficiently-viral videos could swing enough votes to impact a presidential election. Especially wh…

Plausible. I was thinking over the longer-term.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#98
This is good but nowhere as good as EMO https://humanaigc.github.io/emote-portrait-alive/ (https://news.ycombinator.com/item?id=39533326)

This one has too much fake looking body movement and looks eerie/robotic/uncanny valley. The lips don't sync properly in many places. Eye movement and over all head and body movement is not very natural at all.

While EMO looks just perfect mostly. The very first two videos on EMO page are perfect example of that. See the rap near the end to see how good EMO is at lip sync.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#99
post #69

And it's only going to get faster, better, easier, cheaper.[a] Meanwhile, yesterday my credit card company asked me if I wanted to use voice authentication for verifying my identity "more securely" on the phone. Surely the company spent many millions of dollars to enable this new security-theater feature. It begs the question: Is every single executive and manager at my credit card company completely unaware that rig…

> Is every single executive and manager at my credit card company completely unaware that right now anyone can clone anyone else's voice by obtaining a short sample audio clip taken from any social network? Your mistake is assuming the company cares. The "company" is a hundred different disjointed departments that only care about not getting caught Equifax-style (or filing for bankruptcy if caught). If the marketing…

Also, this feature is probably just some midd level execs plan for a bonus, not a rigorously reviewed and planned. It's also probably in the pipeline for a decade so if they don't push it out, suddenly they get no bonus for cancelling a project.

Corporations are ultimately no better than governments and likely worse depending on what their regulatory environment looks like.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#100

This is good but nowhere as good as EMO https://humanaigc.github.io/emote-portrait-alive/ ( https://news.ycombinator.com/item?id=39533326 ) This one has too much fake looking body movement and looks eerie/robotic/uncanny valley. The lips don't sync properly in many places. Eye movement and over all head and body movement is not very natural at all. While EMO looks just perfect mostly. The very first two videos on EMO…

Another research project with 0 model release
Post reply on HN