Live data from Hacker News

Nvidia Uses AI to Slash Bandwidth on Video Calls

petapixel.com

41–50 of 197 posts

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#41
Fundamentally, I don't know if people realise that what we're on the verge of here.

It's effectively a motion-mapped keypoints of the person projected onto a simulated model. I'm assuming the cartoonish avatar was used as an example to partly avoid drawing direct lines to the full implications.

- There's no reason this couldn't extend to voice modelling as well. (much clearer speaking at much lower bandwidth)

- There's no reason this couldn't extend to replacing your sent projection with another image (or person)

- Professional looking suit wearing presentation when you're nude/hungover/unshaven. Hell, why even stop at using your real gender or visage? Imagine a job interview where every candidate, by definition, visually looked the same :)

- There's no reason you couldn't replace other people's avatar with one's of your own choosing as well.

- Why couldn't we model the rest of the environment?

Not there today, but this future is closer than many realise.

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#42

> they have managed to reduce the required bandwidth for a video call by an order of magnitude. In one example, the required data rate fell from 97.28 KB/frame to a measly 0.1165 KB/frame – a reduction to 0.1% of required bandwidth. A nitpick, perhaps, but isn't that three orders of magnitude? We've already seen people use outlandish backgrounds in calls, now it's going to be possible to design similar outlandish vie…

> A nitpick, perhaps, but isn't that three orders of magnitude? I dunno that I'd call it three orders - it's close at about 830x - but it's definitely not even close to being one order either.

Since orders of magnitude are multiplicative, rounded to the nearest whole order of magnitude ~300× to ~3000× is three orders of magnitude.

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#43

Now the person you are speaking to is going to be n% (partially) emulated. n is going to increase in future. One day there will be a paid feature letting you emulate 100% of yourself to respond to video calls when you are not available. And finally, they will replace yourself even without you knowing, and even after you die.

That won't be possible unless full-brain emulation with live data sync gets invented.

Knowledge workers are all about the knowledge that they build up over time when working with a particular environment (be it an industry, a system, or a person/group of people). That knowledge is non-deterministically synthesised in the brain based on the experiences of that person, and being non-deterministic, no AI will come to the same conclusions about every item as this particular human would.

In that case, an emulated personality that is meant to make themselves available as your replacement will be an impostor. One that is less of an expert than yourself at best, and at worst one that is misinformed or misled on various issues (which in turn causes other people to be misled or misinformed).

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#44

Fundamentally, I don't know if people realise that what we're on the verge of here. It's effectively a motion-mapped keypoints of the person projected onto a simulated model. I'm assuming the cartoonish avatar was used as an example to partly avoid drawing direct lines to the full implications. - There's no reason this couldn't extend to voice modelling as well. (much clearer speaking at much lower bandwidth) - There…

[deleted]

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#46
post #19

Watching the video it no longer feels like you're looking at a real person, but instead just another npc. It no longer feels as personal. The last thing remote relationships need is more impersonality. I hope this is used only when it's needed.

I'm really confused. I just watched the video again, and I cannot really see the effect that you're talking about. In fact, I struggle to see any difference from the original. It totally feels genuine, I'm not sure why you perceive them to be npcs

I want to see if people notice this effect when they don't know it's artificially generated. I have a feeling that the uncanniness is at least part because you know it's fake.

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#48

> they have managed to reduce the required bandwidth for a video call by an order of magnitude. In one example, the required data rate fell from 97.28 KB/frame to a measly 0.1165 KB/frame – a reduction to 0.1% of required bandwidth. A nitpick, perhaps, but isn't that three orders of magnitude? We've already seen people use outlandish backgrounds in calls, now it's going to be possible to design similar outlandish vie…

You just need to connect GPT3, and the dialogue is taken care of. Lyrebird’s API will take care of the speech synthesis.

Viola! My deep fake can stand in at meetings now while I code.

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#49

I wonder how weird it gets when you turn your head too much. This is very cool though - I was expecting to be able to tell a difference and maybe slip into uncanny valley territory but it looks good. Big question though - is this just substituting the problem of not having good internet with not having a really fast nVidia graphics card?

If I understand correctly the sender just uses classic object detection/tracking. So the question would be how bad does it look if the receiver just tried to distort the image using that tracking data without having a trained model to smooth out the distortions.

Re: Nvidia Uses AI to Slash Bandwidth on Video Calls

#50
post #21

A technology very similar to this plays a plot point in Vernor Vinge's 1992 novel A Fire Upon the Deep. In his universe, both the interstellar net and combat links between ships are low bandwidth. Hence, video is interpolated between sync frames or recreated from old footage. Vinge calls the resulting video "evocations".

I was thinking of exactly this when I read the article.

The plot point being that when the bandwidth gets too low, the interpolation AI has to make lots of stuff up, you are not quite sure exactly what was said.

I seem to remember the bandwidth in the book was very tiny, small number of bits per second (?) so the AI was taking the speech and compressing it into something more compressed than text then decompressing it at the other end into something that was more or less the same.

Post reply on HN