This is pretty cool (although, I have no idea what other technologies exist for this kind of thing), but it's definitely not convincing enough to a human listener. This sounds like it might be convincing enough for some programs like "Hey, Siri" but it's not gonna convince your mom. You can listen to the samples on the page linked here and you can immediately tell that Obama and Trump don't sound quite human.
Well, the question is, do they just need to throw more computational power / training at this algorithm or is that the peak of their implementation? This is something Google has been working a lot on [1] and Baidu also recently posted about their results too [2]. We're definitely pretty close to passing the human detectable level. [1] https://deepmind.com/blog/wavenet-generative-model-raw-audio... [2] http://research…
The human mind seems to be better at this than most creators credit it for.