Live data from Hacker News

Amazon Polly – Lifelike Text-To-Speech

aws.amazon.com

71–80 of 106 posts

Re: Amazon Polly – Lifelike Text-To-Speech

#71

Earlier quoted context omitted.

That's great to hear! I'm from Gdańsk too :) Out of curiosity - any more info about the team behind Amazon Polly?

Probably the Ivona team, no? https://www.ivona.com/

Yeah IVONA. Sounds similar and after acquisition Amazon was looking for engineers to build cloud offering :-).

Re: Amazon Polly – Lifelike Text-To-Speech

#73
post #10

I'm a native Portuguese speaker and I can say that this is far from how people speak Portuguese. It sounds like pretty much other TTS solutions. However, Spanish and English are awesome and really sound like a person. This is great! Edit: After playing a little bit with it, it does sound better than anything I have seen before even in Portuguese. I believe that the weirdness is just for the standard/demo phrase. When…

Also the sample text in Portuguese has a bad construction

Re: Amazon Polly – Lifelike Text-To-Speech

#74
post #13

One thing I always still notice about these "lifelike" speech models is that they still have random pitch variation that wouldn't be present from a real speaker—sort of a "warbling", similar to audio heavily compressed by some cellular realtime audio codecs. I assume this is due to random variation in the pitches the speakers in the training data used to say given phonemes. Would this, then, imply that speech models…

Additionally when we speak we don't just 'replay' words; we add emotions and so on. Reminds me of this video [1] analysing how the actor Anthony Hopkins converts his lines to speech for his scenes. [1] https://www.youtube.com/watch?v=4kSGkGKwp9U

I guess you could do something like vocaloid does to better "shape" the sounds of each word. I guess its pretty tough to do that automatically though, since words by themselves don't convey the emotions that they should be portrayed with.

I mean, a word or phrase could very much change meaning depending on how its spoken (excitement, disgust, fear, sarcasm etc) and even if the context is available in the text (as is often the case), its rather difficult to machine-detect, at least, for now.

Re: Amazon Polly – Lifelike Text-To-Speech

#75

Earlier quoted context omitted.

The male english sample sounds like dragon speech from 10+ years ago, and the female sounded more robotic than others I've heard... I don't see (or hear?) the advance here...

i thought you guys were being either hyperbolic or overly critical. then i heard the demo files.

98 upvotes? Huh?

Re: Amazon Polly – Lifelike Text-To-Speech

#76
post #38
post #20

Earlier quoted context omitted.

I'm not sure about Mandarin, but this is definitely true for Japanese. Near-human Japanese text-to-speech has been around a very long time.

So good they can even sing, and so well that they create entirely virtual pop stars: https://en.wikipedia.org/wiki/Vocaloid

Vocaloid still sounds robotic (at least, to me). It just so happens when you make autotune/melodyne'd pop songs, it really doesn't matter.

Re: Amazon Polly – Lifelike Text-To-Speech

#77
In our team nobody speaks English as first language. Now we have to hire actors for English voice narration in our demo videos. If Amazon Polly gets better, we might be able to switch to synthesized voice narration and don't deal with actors anymore.

PS.That would be one more job lost to AI.

Re: Amazon Polly – Lifelike Text-To-Speech

#78

Could we drop “Lifelike” from the title? It's marketing puffery (and untrue, it sounded very robotic to me to the point I was convinced I'd heard the wrong audio file). The actual story is that it's Amazon's text-to-speech-as-a-service offering, not that it's particularly innovative in terms of sounding good.

Sooo cynical. This is way better than most TTS systems, including Google's. The only one I've ever heard that is better is DeepWave, and that suffers from lots of background noise.

Seriously? Compare Polly's English samples (particularly #2) in the original article with the Google's WaveNet samples:

Polly 1: https://d0.awsstatic.com/product-marketing/Polly/HelloEnglis...

Polly 2: https://d0.awsstatic.com/product-marketing/Polly/HelloEnglis...

WaveNet 1: https://storage.googleapis.com/deepmind-media/pixie/us-engli...

WaveNet 2: https://storage.googleapis.com/deepmind-media/pixie/us-engli...

There just isn't a comparison -- Polly sounds way more robotic.

Re: Amazon Polly – Lifelike Text-To-Speech

#80

Does anyone know why there is a 1000-character limit?

It says

> The size of the input text can be up to 1500 billed characters (3000 total characters). SSML tags are not counted as billed characters. [1]

Not sure where you got 1000 from. As to the why - no. It doesn't seem like too much. 2 minutes of speech maybe in there?

[1]http://docs.aws.amazon.com/polly/latest/dg/limits.html

Post reply on HN