Live data from Hacker News

Nvidia Uses AI to Slash Bandwidth on Video Calls

petapixel.com

171–180 of 197 posts

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#171
post #137

Earlier quoted context omitted.

In some ways, of course, I agree, but as some other commenters have pointed out, this is essentially a "deepfake" of yourself...in principle there's no restriction on what one could use as their "avatar." As much was demonstrated in the "Different Output Styles" slide. So I'd have a hard time saying this is not dramatically less "real" than some lossy compression technique. Is there some way to formalize "realness?"…

Maybe simply using an objective video quality measurement against the original picture would do the trick. Something like PSNR or SSIM. A "deepfake" is likely to score low on such a score that doesn't depend on high level visual perception. Also know that you can "deepfake" yourself using a traditional video encoder, just change the keyframe to someone else's face. Of course, it will look broken and totally unconvinc…

I think the point is that for very low bandwidth, the "deepfake" version will score higher than the heavily compressed video stream for PSNR.

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#172
post #21

A technology very similar to this plays a plot point in Vernor Vinge's 1992 novel A Fire Upon the Deep. In his universe, both the interstellar net and combat links between ships are low bandwidth. Hence, video is interpolated between sync frames or recreated from old footage. Vinge calls the resulting video "evocations".

I highly recommend A Fire Upon the Deep. Its a rare mix of really interesting hard scifi with an actually good story and characters. Hard scifi often has very flat characters but this is not a book which suffers from it.

It has a very very cool twist to explain the Fermi Paradox and is a really good example of a universe with one modified rule.

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#173

Earlier quoted context omitted.

I posit that no productivity would be lost. Just have a summarizing AI email you the outcomes.

Hello, isoprophlex, I'm your Assistant AI Today's meeting was an hour-and-a-half spent on bike-shedding the position--and color--of the "logout" button on our product page. 15 minutes were spent debating the resident usability expert who suggested white text on a dark blue background would be more readable for people with low vision. The department manager insisted on retaining pastel blue as it is his favorite color…

Next version would be:

45 minutes were spent by humans arguing over whether "logout" or "log out", while I created both buttons (GPT-3 can already do that) and A/B tested it.

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#174
post #85

I would imagine Apple doing this with FaceTime soon. Using their own NPU ( Neural processing unit ), you can now make FaceTime call with ridiculously low bandwidth. From the Nvidia example, 0.1165 KB/frame even at buttery smooth 60fps ( I could literally hear Apple market the crap out of this ), that is 7KBps or 56Kbps! Remember when the industry were trying to compress CD Audio quality ( aka 128Kbps MP3 ) down to 64…

The problem is with phone the background change more, and you show more things than just a static face, because it is easier to move around.

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#175
I see a lot of people being alienated by the fact that people could take on different avatars during their meeting. I would honestly accept that with no question.

In a work environment, I would expect the person I'm talking to to be presentable, ie their avatar would be presentable, so no goofy backgrounds or annoying accessories.

But the key for me is, I'd actually have something to see. So often in my work in in meetings and three people have cameras on and the rest don't. I don't really care what they look like, I care if they're engaged, nodding their heads, their facial reactions.

I don't always have my video on either, I don't have great upload speeds so I usually appear as a big blob anyway. I'd happily have whatever representation of me be in my place if it meant people could see my reactions

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#176

> they have managed to reduce the required bandwidth for a video call by an order of magnitude. In one example, the required data rate fell from 97.28 KB/frame to a measly 0.1165 KB/frame – a reduction to 0.1% of required bandwidth. A nitpick, perhaps, but isn't that three orders of magnitude? We've already seen people use outlandish backgrounds in calls, now it's going to be possible to design similar outlandish vie…

David Foster Wallace predicted this in his novel Infinite Jest. Except they where static images inserted over a video phone, and the user had to keep their head positioned just right to make them work.

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#177

I see a lot of people being alienated by the fact that people could take on different avatars during their meeting. I would honestly accept that with no question. In a work environment, I would expect the person I'm talking to to be presentable, ie their avatar would be presentable, so no goofy backgrounds or annoying accessories. But the key for me is, I'd actually have something to see. So often in my work in in me…

This is why I keep my camera on, so that whoever is talking sees me nodding away whilst they talk and when they go "any questions" you can pause and say "None from me". It's also why you always unmute yourself when someone joins the call - yes, we can hear you, yes your mic is fine, no you're not muted, it's just no one is replying becuase they're all muted and causedoing something else because we haven't started yet.

It feels like people really haven't put any thought into how to handle video conferences.

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#178

> they have managed to reduce the required bandwidth for a video call by an order of magnitude. In one example, the required data rate fell from 97.28 KB/frame to a measly 0.1165 KB/frame – a reduction to 0.1% of required bandwidth. A nitpick, perhaps, but isn't that three orders of magnitude? We've already seen people use outlandish backgrounds in calls, now it's going to be possible to design similar outlandish vie…

You just need to connect GPT3, and the dialogue is taken care of. Lyrebird’s API will take care of the speech synthesis. Viola! My deep fake can stand in at meetings now while I code.

But who's playing the viola?

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#179

> they have managed to reduce the required bandwidth for a video call by an order of magnitude. In one example, the required data rate fell from 97.28 KB/frame to a measly 0.1165 KB/frame – a reduction to 0.1% of required bandwidth. A nitpick, perhaps, but isn't that three orders of magnitude? We've already seen people use outlandish backgrounds in calls, now it's going to be possible to design similar outlandish vie…

I think the ability to, as someone mentioned it, have yourself look a bit tidier than you actually are (working from home) could be a huge benifit. I mean taking away focus on things that doesn't matter in a virtual meeting such as: Where you are sitting - via Virtual Background Your daily hair style status or if you have a nose pimple - Via NVIDIAs AI showcased here. Would be great. Though replacing yourself with a…

It depends on how accurately the avatar is able to represent important information: emotion, attention, state of mind, etc. There's a lot of bandwidth in looking at a face (and bodylanguage as well), that's where the value in face-to-face meetings is.

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#180
post #110

Earlier quoted context omitted.

Video is not real either. The perspective is different and the colors are wrong. It is even less real after it went through lossy compression. The whole point of a lossy compression is to remove details that you perceive as unimportant. For example, leaves on a tree may look like a greeny mess, but that's fine, from afar, you don't make a difference. Using neural networks for compression is far from being a new conce…

The distinction that makes ai compression creepy is that it can use prior knowledge of what a face is. Traditional compression has no high level knowledge of what a video call looks like, the algorithms are about patterns in how pixels change in time and space, so the artifacts from the algorithm are pixel-based effects (blockiness, blurriness, etc). A neural net that has been trained on a million faces and is setup…

You're sitting at your computer, you're watching your colleagues are discussing something in detail, slowly you start to notice - your two black colleagues are starting to literally look the same. That's weird, are you slowly getting more racist? oh right, your bandwith dropped and your AI upscaler has decided to upscale all your black colleagues up to some stereotype it got trained with.

I'm afraid I suspect this isn't that far fetched.

Post reply on HN