A few Superbowls ago Intel demoed a supercomputer which could take in live video streams from hundreds of cameras, run a batch job for about 30 seconds, and then be able to synthesize video from arbitrary viewpoints above the crowd.

I think it worked by making a point cloud that fits the camera observations.

Obstruction is not a big issue in that use case, but if there was obstruction, the system could choose to ignore pixels that were obstructed when constructing the image.