Earlier quoted context omitted.
It looks like they're doing photogrammetry in real time, which is mind boggling. I'm not familiar with this space, but building a 3D model, texturing it with live video, compressing and sending that over the internet, and doing it with minimal latency for it to be believable/enjoyable? Incredible technical achievement if that's the approach. Using state of the art tech, no doubt, and probably lots of ML magic to smoo…
not that hard to do if you have actual depth sensing cameras, and even without those, something like the oculus quest 2 does that exact task (generate a rough 3d volume based on several 2d video feeds) you can see a neat example when you draw your guardian space, and move objects (and notice how it updates the 3d volume representation)
Then it would be done already.