Live data from Hacker News

Amazon Polly – Lifelike Text-To-Speech

aws.amazon.com

31–40 of 106 posts

Re: Amazon Polly – Lifelike Text-To-Speech

#33
post #8

#offtopic a bit but i wonder if anyone knows a good api for the other-way around -> speech to text

Shameless plug: check out Tropo[1]: you can interface with SIP.

One thing to keep in mind regarding speech to text: if you need accuracy, it helps immensely if you can guide the recognizer with a list of expected answers. Un-hinted transcription will do its best to recognize words but without AI to understand the context, there are still quite a few possibilities for mis-recognition. More info about these two here[2].

[1] https://www.tropo.com

[2] https://www.tropo.com/docs/voice/transcription-vs-speech-rec...

(Full disclaimer: I'm in the Tropo BU)

Re: Amazon Polly – Lifelike Text-To-Speech

#35

Could we drop “Lifelike” from the title? It's marketing puffery (and untrue, it sounded very robotic to me to the point I was convinced I'd heard the wrong audio file). The actual story is that it's Amazon's text-to-speech-as-a-service offering, not that it's particularly innovative in terms of sounding good.

The male english sample sounds like dragon speech from 10+ years ago, and the female sounded more robotic than others I've heard... I don't see (or hear?) the advance here...

i thought you guys were being either hyperbolic or overly critical. then i heard the demo files.

Re: Amazon Polly – Lifelike Text-To-Speech

#36

Could we drop “Lifelike” from the title? It's marketing puffery (and untrue, it sounded very robotic to me to the point I was convinced I'd heard the wrong audio file). The actual story is that it's Amazon's text-to-speech-as-a-service offering, not that it's particularly innovative in terms of sounding good.

I agree that the voice sounds pretty bad, but I think the 'lifelike' part comes from the intelligence part. I.E WA -> Washington, and 75F -> 75 Fahrenheit

Re: Amazon Polly – Lifelike Text-To-Speech

#38
post #20
post #13

One thing I always still notice about these "lifelike" speech models is that they still have random pitch variation that wouldn't be present from a real speaker—sort of a "warbling", similar to audio heavily compressed by some cellular realtime audio codecs. I assume this is due to random variation in the pitches the speakers in the training data used to say given phonemes. Would this, then, imply that speech models…

I'm not sure about Mandarin, but this is definitely true for Japanese. Near-human Japanese text-to-speech has been around a very long time.

So good they can even sing, and so well that they create entirely virtual pop stars: https://en.wikipedia.org/wiki/Vocaloid

Re: Amazon Polly – Lifelike Text-To-Speech

#39

Could we drop “Lifelike” from the title? It's marketing puffery (and untrue, it sounded very robotic to me to the point I was convinced I'd heard the wrong audio file). The actual story is that it's Amazon's text-to-speech-as-a-service offering, not that it's particularly innovative in terms of sounding good.

[deleted]

Re: Amazon Polly – Lifelike Text-To-Speech

#40
post #36

Could we drop “Lifelike” from the title? It's marketing puffery (and untrue, it sounded very robotic to me to the point I was convinced I'd heard the wrong audio file). The actual story is that it's Amazon's text-to-speech-as-a-service offering, not that it's particularly innovative in terms of sounding good.

I agree that the voice sounds pretty bad, but I think the 'lifelike' part comes from the intelligence part. I.E WA -> Washington, and 75F -> 75 Fahrenheit

Text-to-speech systems have been expanding abbreviations for decades, and it always backfires in some cases. For example, DECtalk would expand "Sun" to "Sunday" regardless of context, with the humorous result that blind people reading tech news in the 90s would often hear about Sunday Microsystems.

Edit: Yes, that example was unfair, because it's from the 90s (and actually, DECtalk was largely unchanged since the 80s, until it was ruined in the late 90s). Here's a somewhat more recent one: With the ETI-Eloquence engine, which was last updated in 2002, the string "2 Marketing" (which one might find, say, in a Wikipedia table of contents) is expanded to "March twond keting".

Post reply on HN