Live data from Hacker News

VASA-1: Lifelike audio-driven talking faces generated in real time

microsoft.com

151–160 of 166 posts

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#151

What this is starting to reveal is that there's a clear need for some kind of chain of custody system that guarantees the authenticity of what we see. Nikon/Canon tried doing this in the past, but improper storage of private keys lead to vulnerabilities. As far as I'm aware it's never extended to video either. With modern secure hardware keys it may yet be possible. The difficulty is that any kind of photo/video mani…

None of that works, it's simply theatre. I can just take a (crypto-signed) photo of another photo.

The public block chain would show the chain of custody/ownership, so the photo of a photo would show the the final crypto signature does not belong to the claimed owner.

You are correct that I as a viewer can't just rely on a crypto-signature like a watermark, I'd have to verify the chain of custody, but if I wanted to do that, it is available to do so.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#152

This is good but nowhere as good as EMO https://humanaigc.github.io/emote-portrait-alive/ ( https://news.ycombinator.com/item?id=39533326 ) This one has too much fake looking body movement and looks eerie/robotic/uncanny valley. The lips don't sync properly in many places. Eye movement and over all head and body movement is not very natural at all. While EMO looks just perfect mostly. The very first two videos on EMO…

This is real time!

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#153
post #63

Earlier quoted context omitted.

No, we are merely returning to the pre-photography state of things where a mere printed image is not sufficient evidence for anything.

True, an image, audio clip, or video is not enough evidence to establish truth. We still need a way to establish truth. It's important for security cameras, for politics, and for public figures. Here are some things we could start looking into. * Cameras that sign their output. Yes, this camera caught this video, and it hasn't been modified. This is a must for recordings being used in court evidence IMO. Otherwise fr…

Every image is an NFT?

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#154
post #64

Why is this research being done? Is this some kind of arms race? The only purpose of this technology I can think of is getting spies to abuse others. Am I going to have to do AuthN and AuthZ on every phone call and zoom now?

Because the text for this is only a slight variation of the tech for a broad range of legitimate applications? Because even this precise tech has legitimate use cases? > The only purpose of this technology I can think of is getting spies to abuse others. Can you really not think of any other use cases?

Why don't you get more specific about your claims?

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#155
post #64

Earlier quoted context omitted.

Because the text for this is only a slight variation of the tech for a broad range of legitimate applications? Because even this precise tech has legitimate use cases? > The only purpose of this technology I can think of is getting spies to abuse others. Can you really not think of any other use cases?

Why don't you get more specific about your claims?

Jeez. I dunno. Sometimes I just reach my threshold for the time I'm prepared to spend debating with strangers on the internet.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#156

Why is this research being done? Is this some kind of arms race? The only purpose of this technology I can think of is getting spies to abuse others. Am I going to have to do AuthN and AuthZ on every phone call and zoom now?

> It paves the way for real-time engagements with lifelike avatars that emulate human conversational behaviors.

Teams started rolling out Avatars https://techcommunity.microsoft.com/t5/microsoft-teams-blog/..., this would be a step up. I'm not really a fan but that doesn't mean I can excuse the use case.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#157
post #16

maybe making a webpage with 27 videos isn't the greatest web design idea

It's up to your browser on whether those are actually loaded all at once. E.g. on Chrome Desktop with no data saver modes enabled it buffers the first couple seconds of each video then when you play it grabs the remaining MBs for that. That way you can see the videos as quick as you like but not actually have to load all 27 fully just because you opened the page.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#158
post #69

And it's only going to get faster, better, easier, cheaper.[a] Meanwhile, yesterday my credit card company asked me if I wanted to use voice authentication for verifying my identity "more securely" on the phone. Surely the company spent many millions of dollars to enable this new security-theater feature. It begs the question: Is every single executive and manager at my credit card company completely unaware that rig…

Pretty much. You think they’re smart or with it. They’re just lucky fogies

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#159

This is good but nowhere as good as EMO https://humanaigc.github.io/emote-portrait-alive/ ( https://news.ycombinator.com/item?id=39533326 ) This one has too much fake looking body movement and looks eerie/robotic/uncanny valley. The lips don't sync properly in many places. Eye movement and over all head and body movement is not very natural at all. While EMO looks just perfect mostly. The very first two videos on EMO…

There were some misses with emo too, but Hepburn at the end was amazing.

Re: VASA-1: Lifelike audio-driven talking faces generated in real time

#160
Don't know why they're not releasing it right away.

If they can do it, so can someone else and hiding it makes it worse. If it's widely available people will quickly realise that the talking head on YT spouting racist BS is AI. This process needs to happen faster.

Ofc there will still be people who don't care or understand, but there will always be people who are for example racist and don't care if the affirmation for their beliefs comes from a human or a machine.

Post reply on HN