Live data from Hacker News

One-Shot Free-View Neural Talking-Head Synthesis for Video Conferencing

nvidia-research-mingyuliu.com

11–20 of 52 posts

Re: One-Shot Free-View Neural Talking-Head Synthesis for Video Conferencing

#11
post #3

Wow, this works pretty well. Makes me think of that chapter in Infinite Jest where videoconferencing gets popular, until people start using "optimized" computer-rendered images instead of showing their actual faces, at which point everyone goes back to audio-only.

Permutation City has another take, where people virtually meet one another but use masks to hide their emotions.

Re: One-Shot Free-View Neural Talking-Head Synthesis for Video Conferencing

#12
post #5
post #3

Wow, this works pretty well. Makes me think of that chapter in Infinite Jest where videoconferencing gets popular, until people start using "optimized" computer-rendered images instead of showing their actual faces, at which point everyone goes back to audio-only.

I wouldn't mind video conferencing with computer generated avatars. I don't video conference to know what the other person looks like, and in fact knowing what they look like just creates lots of unnecessary bias. I do it for the cues from their gestures, facial expressions, the direction they are looking, etc. With a good tracking setup that works perfectly well today with digital avatars.

I like faces. Why deny a key part of being human?

Re: One-Shot Free-View Neural Talking-Head Synthesis for Video Conferencing

#13
post #3

Wow, this works pretty well. Makes me think of that chapter in Infinite Jest where videoconferencing gets popular, until people start using "optimized" computer-rendered images instead of showing their actual faces, at which point everyone goes back to audio-only.

> at which point everyone goes back to audio-only

That defeats the entire purpose of using facial and body expressions that only video provides.

We already have video filters that remove wrinkles and blemishes in videoconferencing to make you look better.

Even if we replace ourselves entirely with computer-rendered images, they're still going to be reproducing our expressions, movements and gestures, which is what matters.

Re: One-Shot Free-View Neural Talking-Head Synthesis for Video Conferencing

#14
post #5
post #3

Wow, this works pretty well. Makes me think of that chapter in Infinite Jest where videoconferencing gets popular, until people start using "optimized" computer-rendered images instead of showing their actual faces, at which point everyone goes back to audio-only.

I wouldn't mind video conferencing with computer generated avatars. I don't video conference to know what the other person looks like, and in fact knowing what they look like just creates lots of unnecessary bias. I do it for the cues from their gestures, facial expressions, the direction they are looking, etc. With a good tracking setup that works perfectly well today with digital avatars.

Ah yes, the solution to bias: Hide everyone's faces. That'll teach everyone to celebrate our differences.

Re: One-Shot Free-View Neural Talking-Head Synthesis for Video Conferencing

#16
post #7

This makes a request to server to get the result back. Hacker News hug of death has already happened. I wish this was deployable to browsers so it was fully stand alone.

https://www.youtube.com/watch?v=xzLHZbBvKNQ talks of the https://developer.nvidia.com/maxine which requires GPU, so I'm assuming an RTX level GPU will be required for full motion? The demo calls an API on what appears to be a static IP which returns the end result using the algorithm.

The full paper is on https://nvlabs.github.io/face-vid2vid/main.pdf . (It only mentions GPU once, for the training set.)

I'm quite impressed by how NVIDIA Broadcast cleans up a simple webcam image already, on a 3070 GPU; the background blur will get the gap between headphone bridge and head with sharp cuts - it's impressive enough in my books to warrant such a gaming grade GPU for work purposes, if a remote worker.

I have my cam off to the side; I'm really looking forward to being able to try the angle correction!

Re: One-Shot Free-View Neural Talking-Head Synthesis for Video Conferencing

#18
post #3

Wow, this works pretty well. Makes me think of that chapter in Infinite Jest where videoconferencing gets popular, until people start using "optimized" computer-rendered images instead of showing their actual faces, at which point everyone goes back to audio-only.

My phone already has a video call beautification setting built into the OS at the camera level.

We've been skirting the line for a while.

If I could, right now I absolutely would prefer to be sending a synthesized avatar then the real me - my desktop setup doesn't allow very optimal camera placement with large monitors, but for maximum impact I ideally want to send my face making direct eye contact with the camera.

Re: One-Shot Free-View Neural Talking-Head Synthesis for Video Conferencing

#19
post #12
post #5

Earlier quoted context omitted.

I wouldn't mind video conferencing with computer generated avatars. I don't video conference to know what the other person looks like, and in fact knowing what they look like just creates lots of unnecessary bias. I do it for the cues from their gestures, facial expressions, the direction they are looking, etc. With a good tracking setup that works perfectly well today with digital avatars.

I like faces. Why deny a key part of being human?

If you don't know why one might want this, you might be a white male, or at least white, or at least male, or at least a member of the majority race in your locality.

I've definitely wanted this on a few occasions, to avoid being discriminated.

Post reply on HN