Live data from Hacker News

Deep Learning for Siri’s Voice

machinelearning.apple.com

41–50 of 95 posts

Re: Deep Learning for Siri’s Voice

#41
post #3

A research paper published by Apple? About Siri?! Unheard of! Last time I was at an NLP conference wth Apple employees they wouldn't say anything about how Siri speech worked, despite being very inquisitive about everyone else's publications. Good to see some change.

Also glad to see this. Still curious as to why they wouldn't post it as a research paper on arXiv -- what's the point in reinventing the wheel here? I suppose it's nice for publicity, but would be great if they also played nicely with the ecosystem.

The actual paper is available here: http://www.isca-speech.org/archive/Interspeech_2017/abstract...

Re: Deep Learning for Siri’s Voice

#42

Kinda sad to see that the names of the authors are omitted, although you can infer some of them from the quote: > For more details on the new Siri text-to-speech system, see our published paper “Siri On-Device Deep Learning-Guided Unit Selection Text-to-Speech System” [9] T. Capes, P. Coles, A. Conkie, L. Golipour, A. Hadjitarkhani, Q. Hu, N. Huddleston, M. Hunt, J. Li, M. Neeracher, K. Prahallad, T. Raitio, R. Rasip…

Because then it wouldn't be an Apple™ iNovation™.

Let's be honest, these people's names aren't displayed prominently for precisely the same reason as early Atari game developers names weren't.

Re: Deep Learning for Siri’s Voice

#43
post #8

It might seem silly, but I'm looking forward to the first AI talk therapist. Most of the benefit of therapy is the talking, so it's not as crazy as it sounds.

> It might seem silly, but I'm looking forward to the first AI talk therapist. Most of the benefit of therapy is the talking, so it's not as crazy as it sounds. Not crazy at all. At least some therapies provide benefits even with simple non-AI processes: "A meta-analyses of 15 studies, published in this month’s volume of Administration and Policy in Mental Health and Mental Health Services Research, found no signific…

The one grey area though is that when the patient actually does the self-help program. Adherence is a big problem. Meaning a “self-helper” needs to be more disciplined because they don’t have the same pressure/accountability they might have with an actual therapist.

Re: Deep Learning for Siri’s Voice

#44
I couldn't read the paper yet, and also I know very little about this, but listening to the audio samples it seems that one of the most notable changes was the intonation in changing phrases. Did anyone else catch something like that? I'm not sure I'm doing a good job at explaining. If you listen to all iOS11 samples it'll stand out.

Anyway, it's the only way I can still identify this as a fake voice. The intonation always follows the same cadence (not sure if that's the word?). We really shouldn't have overused the word awesome before this kind of thing came along.

There's also a kind of dread too, tbh, this kind of seamless TTS has the potential to change a lot of things. First of all criminals are going to love this, youtube pranksters too. Eventually this will shake up the voice acting industry in a possibly not healthy way for the voice actors, while at the same time allowing projects with a shorter budget to have incredible voice work (also dubbing).

What I think is really important, tho, is that as we move away from the uncanny valley we change our relationships with those voices, our brains don't have the capacity to listen to a voice this real and not imagine it as a person, even for adults.

Ironically at this moment I'm using an old threadless sweatshirt that says "this was supposed to be the future" but nowadays I can honestly say we're getting there.

Re: Deep Learning for Siri’s Voice

#46
post #36

The prosody and and continuity of the speech is dramatically improved. This is hard to do and very impressive (especially given that it is being done on-device). Personally, I'm less pleased with the actual new voice itself, although that is more a subjective judgment. After listening to many hundreds of voice talent auditions for Alexa, it's hard to step back from that level of pickiness.

Story time?

Re: Deep Learning for Siri’s Voice

#47
post #23
post #4

My favorite part is that the runtime runs on device. I moved back to Android, but persistently one thing Apple does that I like is they don't move things to the internet as often as Google does. On Android, you get degraded TTS if the internet is shoddy.

It's two different philosophies. With Apple it's about providing sufficient value such that the consumer will pay a premium for the product. With Google it's about providing the minimum viable value such that the user will provide as much of their data as possible.

Apple also cares strongly about privacy, so there's a lot of stuff they refuse to do in the cloud.

Google also cares about privacy, but only in reverse. They don't want you to have any ;)

Re: Deep Learning for Siri’s Voice

#48
post #4

My favorite part is that the runtime runs on device. I moved back to Android, but persistently one thing Apple does that I like is they don't move things to the internet as often as Google does. On Android, you get degraded TTS if the internet is shoddy.

That makes sense. I always wondered why improvements to Siri's voice required updating the device. I figure it had to be running there, just getting the text from the web service and not the audio.

Re: Deep Learning for Siri’s Voice

#49
post #2

The difference between the Siri voices from iOS 9-11 is startling. I can still here some issues especially at the ends of phrases, but it's extremely good.

I hope someone makes a YouTube video going back to when Siri first launched to show just how much it's evolved.

Listening to those samples I remember how big an advancement iOS 10 felt, but it's nothing compared to 11.

Re: Deep Learning for Siri’s Voice

#50

I couldn't read the paper yet, and also I know very little about this, but listening to the audio samples it seems that one of the most notable changes was the intonation in changing phrases. Did anyone else catch something like that? I'm not sure I'm doing a good job at explaining. If you listen to all iOS11 samples it'll stand out. Anyway, it's the only way I can still identify this as a fake voice. The intonation…

I think you're overstating things. On the one hand, a lot of applications where quality wasn't that critical switched over ages ago. And, on the other hand, any application that would have spent the money on voice acting is still going to pay for both the higher quality and for a sound that isn't the same as everyone else is using. (Note that Siri's new iOS voice is based on a new training set from a new person.)

I do think there are applications that we don't just have today because TTS just isn't good enough. I've had some ideas around Alexa apps related to content that would be TTSd. But the current Polly just isn't human enough. I don't think this is there yet either but it's getting close.

Post reply on HN