Earlier quoted context omitted.
RGB-D based semantic segmentation is certainly a thing. I'm sure it's also been done with video sequences as well.
Yeah I wish the flagship phone manufacturers would put the hardware back into the phone to take 3d photos...even better if you can get point cloud data to go with it. The applications right now are kind of cheesy but they will get better and if the majority of photos taken pivot to including depth information i think it could really drive better capabilities from our phones. Eyes are very hard to make and coordinate,…
Most flagships can do this though, any multicamera phone can get some kind of stereo. Google do it with the PDAF pixels for smart bokeh (they have some nice blog posts about it). I don't know if there is a way to so that in an API though (or to obtain the depth map).
https://ai.googleblog.com/2018/11/learning-to-predict-depth-...