Live data from Hacker News

Deep Learning for Siri’s Voice

machinelearning.apple.com

91–95 of 95 posts

Re: Deep Learning for Siri’s Voice

#91
post #66

I couldn't read the paper yet, and also I know very little about this, but listening to the audio samples it seems that one of the most notable changes was the intonation in changing phrases. Did anyone else catch something like that? I'm not sure I'm doing a good job at explaining. If you listen to all iOS11 samples it'll stand out. Anyway, it's the only way I can still identify this as a fake voice. The intonation…

Regarding voice acting, I think there is something to be said about human expression/ad-lib. Sure, you could generate a natural-sounding voice computer voice, but in the context of arts we’re still a ways to go before a computer can go off script and add just the perfect amount of intonation on a certain word that turns a phrase into an iconic quote. Similarly, we don’t see CGI motion capture replacing Andy Serkis an…

I think this is less likely to hit major films or TV shows, but it will hit the audiobook and video game markets pretty hard.

I'm pretty excited about the video game side.

Re: Deep Learning for Siri’s Voice

#92
post #36

The prosody and and continuity of the speech is dramatically improved. This is hard to do and very impressive (especially given that it is being done on-device). Personally, I'm less pleased with the actual new voice itself, although that is more a subjective judgment. After listening to many hundreds of voice talent auditions for Alexa, it's hard to step back from that level of pickiness.

How'd you get to listen to many hundreds of voice talent audition for Alexa?

Re: Deep Learning for Siri’s Voice

#93
post #81

Earlier quoted context omitted.

> I would assume that the OS X TTS engine is the same as Siri's as it would be available in High Sierra. Nope, macOS uses a different set of voices for TTS, which are named differently as well. They work offline and are nowhere as good as Siri.

I haven't needed to use Siri on my Mac, but presumably that's using the same stuff as iOS?

We're talking about the MacOS TTS engine accessible via the `say` utility and (I believe) accessibility tools, not Siri which is presumably a different software stack and which uses remote processing.

Re: Deep Learning for Siri’s Voice

#94
post #36

The prosody and and continuity of the speech is dramatically improved. This is hard to do and very impressive (especially given that it is being done on-device). Personally, I'm less pleased with the actual new voice itself, although that is more a subjective judgment. After listening to many hundreds of voice talent auditions for Alexa, it's hard to step back from that level of pickiness.

How'd you get to listen to many hundreds of voice talent audition for Alexa?

I led the product management and design team.

Re: Deep Learning for Siri’s Voice

#95
post #27

The iOS 11 Siri sounds like it's a real person talking, it's amazing. Does anyone know if there's an open-source TTS library available with such quality (or if anyone is working on one, from this paper)? I would love to have my home speakers announce things in this voice.

She sounds younger to me, but very natural sounding. Will be interesting to see how Siri on Home Pod works out.

She does sound younger. I think it's because the 's' is more... lispy. She sounds a bit more valley-speak-ish, although that could just be a result of sounding more natural of course.
Post reply on HN