Fundamentally, I don't know if people realise that what we're on the verge of here. It's effectively a motion-mapped keypoints of the person projected onto a simulated model. I'm assuming the cartoonish avatar was used as an example to partly avoid drawing direct lines to the full implications. - There's no reason this couldn't extend to voice modelling as well. (much clearer speaking at much lower bandwidth) - There…
Nvidia Uses AI to Slash Bandwidth on Video Calls
161–170 of 197 posts
Re: Nvidia Uses AI to Slash Bandwidth on Video Calls
#162Re: Nvidia Uses AI to Slash Bandwidth on Video Calls
#163Earlier quoted context omitted.
If the sending machine (that films my face) does the same decoding I know will be on the other end, then diffs the raw video with the decoded video and finally sends both things, then the receiver should be able to always piece together a 100% reproduction of the actual video feed on my end. The transmitter can always predict exactly what the receiver will decode, so the correct amount of data to send is the amount o…
I suspect that this isn't feasible. Otherwise lossy video codecs wouldn't exist (or would be very niche). Because if you sent a perfect diff you would essentially have lossless compression again. So either your idea will revolutionize video compression or the "diff" would bring you back to the ballpark of lossless codecs.
Re: Nvidia Uses AI to Slash Bandwidth on Video Calls
#164In the final form, use "You=" as reference and just press it one at 1 seconds to simulate keyframe.
AMAZING!
Re: Nvidia Uses AI to Slash Bandwidth on Video Calls
#165(Yes, I know this is realtime webcam footage, not recorded footage, I'm just curious).
Re: Nvidia Uses AI to Slash Bandwidth on Video Calls
#166> Next step would be to just predict both sides of the conversation and sever the real-life link entirely.
Gmail already does a little bit of this. Google books appointments over the phone on your behalf.
We're on the road to this...
Re: Nvidia Uses AI to Slash Bandwidth on Video Calls
#167At what resolution? And also, does the output actually resembles the original image? Examples with background other than uniform? Would be nice if they provided more than just screenshots It's not uncommon to see video calls at 100kbs-150kbps, which is ~10KB/s, and this is for 7fps or so, including audio. So "per frame" that would be 1KB or so (more for key frames, less for I frames). So they say it can be 0.1KB, so…
This is probably first-order-model[1] using keyframes. You send only one image each 2 seconds and the mesh of face 30 times per seconds. Then use first-order-model to extrapolate 2 seconds of video from the keyframe. Rinse, repeat. Very doable. AMAZING! The original first-order-model could not do 30 frames per second, but maybe this Nvidia model has some improvements. 1 - https://aliaksandrsiarohin.github.io/first-or…
It's like with those news "the new battery type has been discovered", with very little actual data, just guesswork.
Re: Nvidia Uses AI to Slash Bandwidth on Video Calls
#168Earlier quoted context omitted.
I posit that no productivity would be lost. Just have a summarizing AI email you the outcomes.
Hello, isoprophlex, I'm your Assistant AI Today's meeting was an hour-and-a-half spent on bike-shedding the position--and color--of the "logout" button on our product page. 15 minutes were spent debating the resident usability expert who suggested white text on a dark blue background would be more readable for people with low vision. The department manager insisted on retaining pastel blue as it is his favorite color…
WFH hasn't lessened the amount of bullshit, but it's become more tolerable. I mute my mic and put on a playlist with elevator music.
Now all I need is an AI standin and this summarizer and we've basically achieved universal basic income.
Re: Nvidia Uses AI to Slash Bandwidth on Video Calls
#169Re: Nvidia Uses AI to Slash Bandwidth on Video Calls
#170What I don't like about AI processed images is that they are not real. I can't go past the fact that I am not looking at the picture as it looks like in reality but somehow smart approximation of the world that is not necessarily true.