Example from a news story. After seeing the video, I must admit I’m not sure if I believe the guy or not, which is scary. Edit: Better link from deadmutex below - https://www.youtube.com/watch?v=vqr0oER03SE https://www.wfla.com/8-on-your-side/better-call-behnken/inst...
The video quality is too good. The lighting and movements lack mistakes. It can't be first order model, wav2lip, or any of the relatively new audio to video models.
The audio doesn't suffer from spectral noise, and it matches the lip movements close enough to not be TTS. Voice conversion (VC) introduces pitch issues that are readily apparent, and it's incredibly hard to train VC models without a ton of parallel audio data from source and target speakers.
This is absolutely a lie (not a deepfake) and I'd bet money on it.
[1] I created https://fakeyou.com cartoon and celebrity TTS, real time voice to voice mapping for VTubers, and am currently working on ML blendshapes.