Live data from Hacker News

Amazon Polly – Lifelike Text-To-Speech

aws.amazon.com

21–30 of 106 posts

Re: Amazon Polly – Lifelike Text-To-Speech

#21

Could we drop “Lifelike” from the title? It's marketing puffery (and untrue, it sounded very robotic to me to the point I was convinced I'd heard the wrong audio file). The actual story is that it's Amazon's text-to-speech-as-a-service offering, not that it's particularly innovative in terms of sounding good.

The male english sample sounds like dragon speech from 10+ years ago, and the female sounded more robotic than others I've heard...

I don't see (or hear?) the advance here...

Re: Amazon Polly – Lifelike Text-To-Speech

#22
post #7

Previous discussion: https://news.ycombinator.com/item?id=13072944 . For a service claiming to be "lifelike", Joey has totally phrased that question as a statement.

I agree, it doesn't sound "lifelike" at all to me. Compared to WaveNet[1] it's day and night. [1] https://deepmind.com/blog/wavenet-generative-model-raw-audio...

The "babbling" samples here are really fun. Would make for a great "Sims" language.

Re: Amazon Polly – Lifelike Text-To-Speech

#23
post #7

Previous discussion: https://news.ycombinator.com/item?id=13072944 . For a service claiming to be "lifelike", Joey has totally phrased that question as a statement.

I agree, it doesn't sound "lifelike" at all to me. Compared to WaveNet[1] it's day and night. [1] https://deepmind.com/blog/wavenet-generative-model-raw-audio...

Wow, that was actually impressive. A few of those were hard to distiguish from human speakers. I love how they included breathing and mouth-sounds as well, really brings it to life.

Re: Amazon Polly – Lifelike Text-To-Speech

#25
post #13

One thing I always still notice about these "lifelike" speech models is that they still have random pitch variation that wouldn't be present from a real speaker—sort of a "warbling", similar to audio heavily compressed by some cellular realtime audio codecs. I assume this is due to random variation in the pitches the speakers in the training data used to say given phonemes. Would this, then, imply that speech models…

Additionally when we speak we don't just 'replay' words; we add emotions and so on. Reminds me of this video [1] analysing how the actor Anthony Hopkins converts his lines to speech for his scenes.

[1] https://www.youtube.com/watch?v=4kSGkGKwp9U

Re: Amazon Polly – Lifelike Text-To-Speech

#26
The thing these voices always miss is context. We need about 700 variations of every textual sentence that all depend on the mood or goal of the conversation and modulate pitch, volume, speed, word choice, etc. appropriately.

This experiment does cover the "generic computer voice" role quite well, but customers don't want that on e.g. their website. I would want a "I'm a professional designer" voice. So there will always be markets for more appropriate voices [*]. But this one does sound very nice at what it does.

The same will happen with robots, btw. Once we get a generic humanoid robot that can do everything, we will still seek aesthetic variations and employ remote body actors.

Post reply on HN