Live data from Hacker News

Nvidia Uses AI to Slash Bandwidth on Video Calls

petapixel.com

101–110 of 197 posts

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#101
post #94
post #85

I would imagine Apple doing this with FaceTime soon. Using their own NPU ( Neural processing unit ), you can now make FaceTime call with ridiculously low bandwidth. From the Nvidia example, 0.1165 KB/frame even at buttery smooth 60fps ( I could literally hear Apple market the crap out of this ), that is 7KBps or 56Kbps! Remember when the industry were trying to compress CD Audio quality ( aka 128Kbps MP3 ) down to 64…

I think Face ID could be used to create the point map, instead of generating it from an image. They could also use Face ID to prevent the tech from malicious deep fakes, e.g only allow people to use this feature I when Face ID confirms the user and the manipulated photo are the same person.

Isn’t the point map supposedly encrypted on the Secure Enclave?[a] If it is, you can’t access it from the main CPU; you can only ask it to compare a provided one with the stored one.

[a]: I know the password is stored on it, but idk about the face mesh

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#102

> they have managed to reduce the required bandwidth for a video call by an order of magnitude. In one example, the required data rate fell from 97.28 KB/frame to a measly 0.1165 KB/frame – a reduction to 0.1% of required bandwidth. A nitpick, perhaps, but isn't that three orders of magnitude? We've already seen people use outlandish backgrounds in calls, now it's going to be possible to design similar outlandish vie…

I think you are right on the money with your thoughts on virtual avatars. I've already noticed this phenomenon cropping up in some niches.

1. the phenomenon of VTubers https://en.m.wikipedia.org/wiki/Virtual_YouTuber

2. in the virtual animal crossing late night show, Animal Talking, the presenter's (Gary Whitta) avatar doesn't really resemble how the presenter looks in real life https://en.m.wikipedia.org/wiki/Animal_Talking_with_Gary_Whi...

3. I watch a lot of interview s with people in VR Chat and it's very interesting how people seem to find it easier(?) to open up while they are embodying a character. https://youtu.be/KZWOXgc7PA4

Being able to experiment with identity in this way is really interesting to me, and I hope it becomes more mainstream with the proliferation of this technology

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#103
post #85

I would imagine Apple doing this with FaceTime soon. Using their own NPU ( Neural processing unit ), you can now make FaceTime call with ridiculously low bandwidth. From the Nvidia example, 0.1165 KB/frame even at buttery smooth 60fps ( I could literally hear Apple market the crap out of this ), that is 7KBps or 56Kbps! Remember when the industry were trying to compress CD Audio quality ( aka 128Kbps MP3 ) down to 64…

> Remember when the industry were trying to compress CD Audio quality (aka 128Kbps MP3) down to 64Kbps?

128kb mp3 is good enough for most people most of the time, but it isn't CD quality. Having said that, 64kb Opus is almost or about as good as 128kb mp3.

I wonder how well these techniques can be applied to audio.

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#104
post #51

> they have managed to reduce the required bandwidth for a video call by an order of magnitude. In one example, the required data rate fell from 97.28 KB/frame to a measly 0.1165 KB/frame – a reduction to 0.1% of required bandwidth. A nitpick, perhaps, but isn't that three orders of magnitude? We've already seen people use outlandish backgrounds in calls, now it's going to be possible to design similar outlandish vie…

The avatar thing isn't one-sided either: it'd be an awesome power to have to remake others! Real-time silly hats for people I talk to and I'm sold.

"Visualize your audience naked." they said. "Helps calm the nerves."

They had no idea.

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#105

What I don't like about AI processed images is that they are not real. I can't go past the fact that I am not looking at the picture as it looks like in reality but somehow smart approximation of the world that is not necessarily true.

any lossy video-compression algorithm face the same challenge. what you are seeing is artificial and is constructed by the algorithm to minimize the perceptual difference between the real and the constructed video feed.

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#106
post #52

My first thought was about the diversity of faces used in the demo and how ten years ago, computers didn't think black people were humans. https://www.youtube.com/watch?v=t4DT3tQqgRM But after that, I was reminded of the paranoia (or not?) around Zoom and that, for an extreme example, the CCP was mining and generating facial fingerprints and social networks using video calls. It seems like this technology is the same…

If I had to guess, the issue around "black people" is that photos are 2D. We don't really understand just how little information is actually in a photo (we add huge amounts of info in our perception). My guess is that predictive systems are using contrast as a guide to essentially 3D structures which, simply, just cannot be reconstructed from 2D. And therefore, probably struggle more on dark faces which have differen…

While contrast is almost certainly part of it, I’d hazard a guess that the training set is also partly to blame.

Now, I don’t know much about neural networks (AI), but my understanding is that if you provide a training set representative of the population makeup, (in America at least) it’ll be biased towards white people as it hasn’t “seen” enough black person images. My limited understanding would then make me think one would need equal white person photos as well as black person photos.

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#107
post #85

I would imagine Apple doing this with FaceTime soon. Using their own NPU ( Neural processing unit ), you can now make FaceTime call with ridiculously low bandwidth. From the Nvidia example, 0.1165 KB/frame even at buttery smooth 60fps ( I could literally hear Apple market the crap out of this ), that is 7KBps or 56Kbps! Remember when the industry were trying to compress CD Audio quality ( aka 128Kbps MP3 ) down to 64…

> Remember when the industry were trying to compress CD Audio quality (aka 128Kbps MP3) down to 64Kbps? 128kb mp3 is good enough for most people most of the time, but it isn't CD quality. Having said that, 64kb Opus is almost or about as good as 128kb mp3. I wonder how well these techniques can be applied to audio.

I doubt the general technique described here would have any applicability to audio, unless the idea is to dynamically create a text-to-speech model of the speaker and transmit only text. I don’t see that being practical any time soon.

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#109
post #62

Earlier quoted context omitted.

You just need to connect GPT3, and the dialogue is taken care of. Lyrebird’s API will take care of the speech synthesis. Viola! My deep fake can stand in at meetings now while I code.

Imagine if everyone did that

This is one of the themes brilliantly explored (in my opinion), in the (largely hard) scifi book "Lady of Mazes" by Karl Schroeder.

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#110

What I don't like about AI processed images is that they are not real. I can't go past the fact that I am not looking at the picture as it looks like in reality but somehow smart approximation of the world that is not necessarily true.

Video is not real either. The perspective is different and the colors are wrong.

It is even less real after it went through lossy compression. The whole point of a lossy compression is to remove details that you perceive as unimportant. For example, leaves on a tree may look like a greeny mess, but that's fine, from afar, you don't make a difference.

Using neural networks for compression is far from being a new concept, and the result is not more or less real than any other technique. It is just that Nvidia implementation is really good at keeping the most important details in a small size.

If you want a more "real" image, you can just use the AI as a predictor and use the remaining bits to encode the difference, like in a traditional MPEG-style codec.

Post reply on HN