Earlier quoted context omitted.
I agree that the voice sounds pretty bad, but I think the 'lifelike' part comes from the intelligence part. I.E WA -> Washington, and 75F -> 75 Fahrenheit
Text-to-speech systems have been expanding abbreviations for decades, and it always backfires in some cases. For example, DECtalk would expand "Sun" to "Sunday" regardless of context, with the humorous result that blind people reading tech news in the 90s would often hear about Sunday Microsystems. Edit: Yes, that example was unfair, because it's from the 90s (and actually, DECtalk was largely unchanged since the 80s…
Amazon Polly – Lifelike Text-To-Speech
41–50 of 106 posts
Re: Amazon Polly – Lifelike Text-To-Speech
#42One thing I always still notice about these "lifelike" speech models is that they still have random pitch variation that wouldn't be present from a real speaker—sort of a "warbling", similar to audio heavily compressed by some cellular realtime audio codecs. I assume this is due to random variation in the pitches the speakers in the training data used to say given phonemes. Would this, then, imply that speech models…
Additionally when we speak we don't just 'replay' words; we add emotions and so on. Reminds me of this video [1] analysing how the actor Anthony Hopkins converts his lines to speech for his scenes. [1] https://www.youtube.com/watch?v=4kSGkGKwp9U
Re: Amazon Polly – Lifelike Text-To-Speech
#43Could we drop “Lifelike” from the title? It's marketing puffery (and untrue, it sounded very robotic to me to the point I was convinced I'd heard the wrong audio file). The actual story is that it's Amazon's text-to-speech-as-a-service offering, not that it's particularly innovative in terms of sounding good.
Re: Amazon Polly – Lifelike Text-To-Speech
#44One thing I always still notice about these "lifelike" speech models is that they still have random pitch variation that wouldn't be present from a real speaker—sort of a "warbling", similar to audio heavily compressed by some cellular realtime audio codecs. I assume this is due to random variation in the pitches the speakers in the training data used to say given phonemes. Would this, then, imply that speech models…
Re: Amazon Polly – Lifelike Text-To-Speech
#45One thing I always still notice about these "lifelike" speech models is that they still have random pitch variation that wouldn't be present from a real speaker—sort of a "warbling", similar to audio heavily compressed by some cellular realtime audio codecs. I assume this is due to random variation in the pitches the speakers in the training data used to say given phonemes. Would this, then, imply that speech models…
Sheer speculation, but I wonder if that could be related to the "checkerboarding" artifacts you almost always see in AI-generated images? Here's a good writeup: http://distill.pub/2016/deconv-checkerboard/
I doubt this would result in audible artifacts at the audio level, though. The visual checkerboarding appears at the 1px detail-level of the output image—this would translate more to minor fluctuations along an audio sample's Nyquist frequency, which is effectively undetectable if the output sampling rate is of the usual kind (44.1kHz).
Re: Amazon Polly – Lifelike Text-To-Speech
#46Re: Amazon Polly – Lifelike Text-To-Speech
#47Re: Amazon Polly – Lifelike Text-To-Speech
#48One thing I always still notice about these "lifelike" speech models is that they still have random pitch variation that wouldn't be present from a real speaker—sort of a "warbling", similar to audio heavily compressed by some cellular realtime audio codecs. I assume this is due to random variation in the pitches the speakers in the training data used to say given phonemes. Would this, then, imply that speech models…
I'm not sure about Mandarin, but this is definitely true for Japanese. Near-human Japanese text-to-speech has been around a very long time.
Re: Amazon Polly – Lifelike Text-To-Speech
#49Re: Amazon Polly – Lifelike Text-To-Speech
#50Could we drop “Lifelike” from the title? It's marketing puffery (and untrue, it sounded very robotic to me to the point I was convinced I'd heard the wrong audio file). The actual story is that it's Amazon's text-to-speech-as-a-service offering, not that it's particularly innovative in terms of sounding good.
Sooo cynical. This is way better than most TTS systems, including Google's. The only one I've ever heard that is better is DeepWave, and that suffers from lots of background noise.