Very cool. I shudder to think of the GPU costs to run these models though. Perhaps they're using TPUs to be as efficient as possible. If you imagine a room occupied in the evening by people for several hours and you have any decent framerate, you're running your pose estimation network on each frame for several hours. And these models are big as far as I have seen. So that pretty much means you have one cloud GPU per…
I think you're overestimating the processing requirements. The original 2010 Kinect did fundamentally similar processing (multi-person tracking and skeletal mapping) on the Xbox 360 which had a PowerPC CPU from 2005.
Was the Kinect even a neural network? I don't think it was.