You could then use the camera perspectives to create a 3d image of the person you're conversing with and map the colour data correctly to that 3d image (Photogrammetry)
You could also likely use the information from the four cameras to map the orientation of the 3d image of the person you're speaking with to give you that sense of depth as you shift your position. https://www.youtube.com/watch?v=Jd3-eiid-Uw
If you had a speaker and mic in each corner you could also capture / emit subtle differences in audio to further enhance it.