Live data from Hacker News

Amazon Polly – Lifelike Text-To-Speech

aws.amazon.com

11–20 of 106 posts

Re: Amazon Polly – Lifelike Text-To-Speech

#11
Given that this supports Speech Synthesis Markup Language [0], I wonder what the impact will be on positions that talk for a living. While this is unlikely to replace voice actors for games/cartoons, I wonder if public service announcers and such will be replaced by this robot.

https://www.w3.org/TR/speech-synthesis/

Re: Amazon Polly – Lifelike Text-To-Speech

#12

Congrats to the team behind the release in Gdańsk, Poland!

That's great to hear! I'm from Gdańsk too :) Out of curiosity - any more info about the team behind Amazon Polly?

Probably the Ivona team, no? https://www.ivona.com/

Re: Amazon Polly – Lifelike Text-To-Speech

#13
One thing I always still notice about these "lifelike" speech models is that they still have random pitch variation that wouldn't be present from a real speaker—sort of a "warbling", similar to audio heavily compressed by some cellular realtime audio codecs.

I assume this is due to random variation in the pitches the speakers in the training data used to say given phonemes. Would this, then, imply that speech models for tonal languages like Mandarin don't share this problem, since people are speaking with prescribed pitches more often?

Re: Amazon Polly – Lifelike Text-To-Speech

#15
post #13

One thing I always still notice about these "lifelike" speech models is that they still have random pitch variation that wouldn't be present from a real speaker—sort of a "warbling", similar to audio heavily compressed by some cellular realtime audio codecs. I assume this is due to random variation in the pitches the speakers in the training data used to say given phonemes. Would this, then, imply that speech models…

I was wondering why it sometimes sounded like bad compression artifacts! Thanks.

Re: Amazon Polly – Lifelike Text-To-Speech

#16
Could we drop “Lifelike” from the title? It's marketing puffery (and untrue, it sounded very robotic to me to the point I was convinced I'd heard the wrong audio file). The actual story is that it's Amazon's text-to-speech-as-a-service offering, not that it's particularly innovative in terms of sounding good.

Re: Amazon Polly – Lifelike Text-To-Speech

#18
post #13

One thing I always still notice about these "lifelike" speech models is that they still have random pitch variation that wouldn't be present from a real speaker—sort of a "warbling", similar to audio heavily compressed by some cellular realtime audio codecs. I assume this is due to random variation in the pitches the speakers in the training data used to say given phonemes. Would this, then, imply that speech models…

Sheer speculation, but I wonder if that could be related to the "checkerboarding" artifacts you almost always see in AI-generated images?

Here's a good writeup: http://distill.pub/2016/deconv-checkerboard/

Re: Amazon Polly – Lifelike Text-To-Speech

#19
post #13

One thing I always still notice about these "lifelike" speech models is that they still have random pitch variation that wouldn't be present from a real speaker—sort of a "warbling", similar to audio heavily compressed by some cellular realtime audio codecs. I assume this is due to random variation in the pitches the speakers in the training data used to say given phonemes. Would this, then, imply that speech models…

I always wonder why they don't program a second pass over the output with a model, say, trained on radio conversations. That second model just for audio 'quality'.

Re: Amazon Polly – Lifelike Text-To-Speech

#20
post #13

One thing I always still notice about these "lifelike" speech models is that they still have random pitch variation that wouldn't be present from a real speaker—sort of a "warbling", similar to audio heavily compressed by some cellular realtime audio codecs. I assume this is due to random variation in the pitches the speakers in the training data used to say given phonemes. Would this, then, imply that speech models…

I'm not sure about Mandarin, but this is definitely true for Japanese. Near-human Japanese text-to-speech has been around a very long time.
Post reply on HN