Live data from Hacker News

Nvidia Uses AI to Slash Bandwidth on Video Calls

petapixel.com

191–197 of 197 posts

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#191

Earlier quoted context omitted.

If the sending machine (that films my face) does the same decoding I know will be on the other end, then diffs the raw video with the decoded video and finally sends both things, then the receiver should be able to always piece together a 100% reproduction of the actual video feed on my end. The transmitter can always predict exactly what the receiver will decode, so the correct amount of data to send is the amount o…

You’re thinking along the right lines, but the challenge is that a raw diff will have the same number of pixels as the raw image, so no compression in bandwidth. So, how do we represent the diff/residue also with fewer numbers? At that point it’s the same as choosing better parameters within some clever encoding (be it pre-designed like JPEG or H.264 or learned via ML).

I was thinking that subtracting the predicted image would give an image that has more zeroes and compresses better (much like dct+quantization for jpeg). After all, any time the neural network would predict an area of the image almost exactly, it can be omitted from the diff stream completely too.

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#192
post #40

Earlier quoted context omitted.

I think it is three orders of magnitude. "An order-of-magnitude difference between two values is a factor of 10. For example, the mass of the planet Saturn is 95 times that of Earth, so Saturn is two orders of magnitude more massive than Earth." https://en.wikipedia.org/wiki/Order_of_magnitude

> "More precisely, the order of magnitude of a number can be defined in terms of the common logarithm, usually as the integer part of the logarithm, obtained by truncation." $ bc -l l(835)/l(10) 2.92168647548360208478 That would make it 2 orders of magnitude by that method. Happy to accept that it's 3 orders of magnitude by the N=a*10^b method though. Either way, it's definitely not one.

I think the whole point of orders of magnitude is to be a back-of-the-napkin estimation of what's what.

"a new car is an order of magnitude difference in price compared to a used car" is appropriate even if a new car is 40k and a used car is 5k

Electric cars have two orders of magnitude less energy storage than gasoline cars, but newer ones are only one order of magnitude.

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#193

> they have managed to reduce the required bandwidth for a video call by an order of magnitude. In one example, the required data rate fell from 97.28 KB/frame to a measly 0.1165 KB/frame – a reduction to 0.1% of required bandwidth. A nitpick, perhaps, but isn't that three orders of magnitude? We've already seen people use outlandish backgrounds in calls, now it's going to be possible to design similar outlandish vie…

I think you are right on the money with your thoughts on virtual avatars. I've already noticed this phenomenon cropping up in some niches. 1. the phenomenon of VTubers https://en.m.wikipedia.org/wiki/Virtual_YouTuber 2. in the virtual animal crossing late night show, Animal Talking, the presenter's (Gary Whitta) avatar doesn't really resemble how the presenter looks in real life https://en.m.wikipedia.org/wiki/Animal…

I'm reminded of Permutation City where they talk to one another with virtual avatars that are able to mask emotional responses and such.

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#194

Earlier quoted context omitted.

You’re thinking along the right lines, but the challenge is that a raw diff will have the same number of pixels as the raw image, so no compression in bandwidth. So, how do we represent the diff/residue also with fewer numbers? At that point it’s the same as choosing better parameters within some clever encoding (be it pre-designed like JPEG or H.264 or learned via ML).

I was thinking that subtracting the predicted image would give an image that has more zeroes and compresses better (much like dct+quantization for jpeg). After all, any time the neural network would predict an area of the image almost exactly, it can be omitted from the diff stream completely too.

The diff will be low norm (for some suitable norm), but it needn’t be sparse in pixel space. It could be sparse in some other spaces (Eg: wavelet basis), but that’s right where the challenge lies — finding a basis where the data (and/or the residue) to be transmitted is sparse.

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#196

I see a lot of people being alienated by the fact that people could take on different avatars during their meeting. I would honestly accept that with no question. In a work environment, I would expect the person I'm talking to to be presentable, ie their avatar would be presentable, so no goofy backgrounds or annoying accessories. But the key for me is, I'd actually have something to see. So often in my work in in me…

The camera on or off brings up another question: When do you look straight into the camera and when do you look at the screen to see the person speaking? I can see argument for both.

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#197
post #11

Stills OK, it would be interesting to see it move. Risk for uncanny valley? Petapixel is a blog spam site btw. Why not go to the source that is linked in the post?

There is a video at the top of the article.

See it now, Firefox Klar wasn't willing to play it.
Post reply on HN