Earlier quoted context omitted.
In a generation or two they'll do facial tracking. I imagine you'll get a full 3D model made of you, and it'll pick up the facial movements and transpose them onto your avatar
This is possible with current consumer-grade webcams. You'll video conference like you do today, except you'll see life-like rendered avatars of each participant. Participants with VR headsets will be able to see everyone else sitting at a table with them. If you don't have a headset, you'll see the same video conference UI you see today, just with avatars instead of video.
And never come back.