One thing I always still notice about these "lifelike" speech models is that they still have random pitch variation that wouldn't be present from a real speaker—sort of a "warbling", similar to audio heavily compressed by some cellular realtime audio codecs. I assume this is due to random variation in the pitches the speakers in the training data used to say given phonemes. Would this, then, imply that speech models…
It's worth noting that purely synthetic speech, while lacking the superficial lifelike quality, doesn't have these little inconsistencies. Perhaps that's one reason why many blind people, particularly power users, prefer pure synthetic text-to-speech engines, such as ETI-Eloquence (commercial) and eSpeak (open source). These systems especially sound better at high speeds than the newer "lifelike" ones.
Just for fun, here's a side-by-side comparison between a concatenative system, IVONA (acquired by Amazon):
http://mwcampbell.us/tmp/derefr_comment_ivona.mp3
And a purely synthetic system, Eloquence:
http://mwcampbell.us/tmp/derefr_comment_eloquence.mp3
The latter is the text-to-speech engine I use every day, and that clip was synthesized at the speaking rate I normally use.
Edit: It appears the word I was looking for when I said "purely synthetic" is "parametric".