It's effectively a motion-mapped keypoints of the person projected onto a simulated model. I'm assuming the cartoonish avatar was used as an example to partly avoid drawing direct lines to the full implications.
- There's no reason this couldn't extend to voice modelling as well. (much clearer speaking at much lower bandwidth)
- There's no reason this couldn't extend to replacing your sent projection with another image (or person)
- Professional looking suit wearing presentation when you're nude/hungover/unshaven. Hell, why even stop at using your real gender or visage? Imagine a job interview where every candidate, by definition, visually looked the same :)
- There's no reason you couldn't replace other people's avatar with one's of your own choosing as well.
- Why couldn't we model the rest of the environment?
Not there today, but this future is closer than many realise.