Interesting. They are not TTS like we are accustomed to, they are replicating a specific persons voice with TTS. Listen to the ground-truth recordings at the bottom and then the synthesized versions above. "Fake News" is about to get a lot more compelling when you can make anyone say anything as long as you have some previous recordings of their voice.
Adobe has already developed that technology:
https://arstechnica.co.uk/information-technology/2016/11/ado...
Now imagine combining it with this:
Face2Face: Real-time Face Capture and Reenactment of RGB Videos https://www.youtube.com/watch?v=ohmajJTcpNk
Perhaps using the intonation from the face-actor's voice to guide the speech synthesis.