While all of these vec2speech type models are impressive, I get the feeling that most of the comments didn't listen to any of the samples. It's still distinctly robotic sounding, probably has quite a bit of garbage output that needs to be filtered manually (as many of these nets often have) and is a far cry from fooling a human.
It doesn't sound too different from a voice coming over a walkie talkie or some kind of intercom. The problem might be that high frequencies, especially overtones, aren't properly constructed, but I'm certain that can be improved.
You can synthesize someone's voice perfectly, but if it's stressing words incorrectly or not at all, it's not going to fool anyone.
Then again, that's probably easier to work around by having humans annotate the sentences to be read.