Amazon Polly – Lifelike Text-To-Speech
11–20 of 106 posts
Re: Amazon Polly – Lifelike Text-To-Speech
#12Congrats to the team behind the release in Gdańsk, Poland!
That's great to hear! I'm from Gdańsk too :) Out of curiosity - any more info about the team behind Amazon Polly?
Re: Amazon Polly – Lifelike Text-To-Speech
#13I assume this is due to random variation in the pitches the speakers in the training data used to say given phonemes. Would this, then, imply that speech models for tonal languages like Mandarin don't share this problem, since people are speaking with prescribed pitches more often?
Re: Amazon Polly – Lifelike Text-To-Speech
#14#offtopic a bit but i wonder if anyone knows a good api for the other-way around -> speech to text
Re: Amazon Polly – Lifelike Text-To-Speech
#15One thing I always still notice about these "lifelike" speech models is that they still have random pitch variation that wouldn't be present from a real speaker—sort of a "warbling", similar to audio heavily compressed by some cellular realtime audio codecs. I assume this is due to random variation in the pitches the speakers in the training data used to say given phonemes. Would this, then, imply that speech models…
Re: Amazon Polly – Lifelike Text-To-Speech
#16Re: Amazon Polly – Lifelike Text-To-Speech
#17Re: Amazon Polly – Lifelike Text-To-Speech
#18One thing I always still notice about these "lifelike" speech models is that they still have random pitch variation that wouldn't be present from a real speaker—sort of a "warbling", similar to audio heavily compressed by some cellular realtime audio codecs. I assume this is due to random variation in the pitches the speakers in the training data used to say given phonemes. Would this, then, imply that speech models…
Here's a good writeup: http://distill.pub/2016/deconv-checkerboard/
Re: Amazon Polly – Lifelike Text-To-Speech
#19One thing I always still notice about these "lifelike" speech models is that they still have random pitch variation that wouldn't be present from a real speaker—sort of a "warbling", similar to audio heavily compressed by some cellular realtime audio codecs. I assume this is due to random variation in the pitches the speakers in the training data used to say given phonemes. Would this, then, imply that speech models…
Re: Amazon Polly – Lifelike Text-To-Speech
#20One thing I always still notice about these "lifelike" speech models is that they still have random pitch variation that wouldn't be present from a real speaker—sort of a "warbling", similar to audio heavily compressed by some cellular realtime audio codecs. I assume this is due to random variation in the pitches the speakers in the training data used to say given phonemes. Would this, then, imply that speech models…