Instead of a "deep fake" face swap an attacker could send virtual video from a fully-virtual environment using something like an nvidia Metahuman controlled by the camera array. I think that would be pretty easily detectable today but maybe less so with an emulated bad webcam and low-res video link. The models/rigging are only going to improve in the future.
The classic "Put a shoe on your head" verification route would still defeat that, at least until someone invents a very good tool to allow those types of models to spawn and manipulate props.