Live data from Hacker News

Deep Learning for Siri’s Voice

machinelearning.apple.com

21–30 of 95 posts

Re: Deep Learning for Siri’s Voice

#22
post #8

It might seem silly, but I'm looking forward to the first AI talk therapist. Most of the benefit of therapy is the talking, so it's not as crazy as it sounds.

There are ongoing efforts in this direction, e.g. this paper from the just concluded Interspeech 2017: http://bit.ly/2wBgLKC

Re: Deep Learning for Siri’s Voice

#23
post #4

My favorite part is that the runtime runs on device. I moved back to Android, but persistently one thing Apple does that I like is they don't move things to the internet as often as Google does. On Android, you get degraded TTS if the internet is shoddy.

It's two different philosophies. With Apple it's about providing sufficient value such that the consumer will pay a premium for the product. With Google it's about providing the minimum viable value such that the user will provide as much of their data as possible.

Re: Deep Learning for Siri’s Voice

#24

Earlier quoted context omitted.

Also glad to see this. Still curious as to why they wouldn't post it as a research paper on arXiv -- what's the point in reinventing the wheel here? I suppose it's nice for publicity, but would be great if they also played nicely with the ecosystem.

This was presented at Interspeech '17 this morning. Maybe the paper is embargoed or something until a later date?

Yeah, Interspeech IS as traditional an ecosystem as it gets in this field.

Re: Deep Learning for Siri’s Voice

#25
Kinda sad to see that the names of the authors are omitted, although you can infer some of them from the quote:

> For more details on the new Siri text-to-speech system, see our published paper “Siri On-Device Deep Learning-Guided Unit Selection Text-to-Speech System”

[9] T. Capes, P. Coles, A. Conkie, L. Golipour, A. Hadjitarkhani, Q. Hu, N. Huddleston, M. Hunt, J. Li, M. Neeracher, K. Prahallad, T. Raitio, R. Rasipuram, G. Townsend, B. Williamson, D. Winarsky, Z. Wu, H. Zhang. Siri On-Device Deep Learning-Guided Unit Selection Text-to-Speech System, Interspeech, 2017.

Why not just add the names by default?

Re: Deep Learning for Siri’s Voice

#26
post #3

A research paper published by Apple? About Siri?! Unheard of! Last time I was at an NLP conference wth Apple employees they wouldn't say anything about how Siri speech worked, despite being very inquisitive about everyone else's publications. Good to see some change.

It's probably safe to assume a lot of that was due to some/most of Siri being licensed from Nuance initially. I mean, who wants to talk about a new product, which most people think is brand new and entirely innovative, just to say "Oh yeah, we paid someone else to work with us to create it."

Not that there's anything wrong with that and it certainly seems like Apple has been investing in-house pretty heavily in recent years for Siri improvement.

Re: Deep Learning for Siri’s Voice

#27

The iOS 11 Siri sounds like it's a real person talking, it's amazing. Does anyone know if there's an open-source TTS library available with such quality (or if anyone is working on one, from this paper)? I would love to have my home speakers announce things in this voice.

She sounds younger to me, but very natural sounding.

Will be interesting to see how Siri on Home Pod works out.

Re: Deep Learning for Siri’s Voice

#28

Kinda sad to see that the names of the authors are omitted, although you can infer some of them from the quote: > For more details on the new Siri text-to-speech system, see our published paper “Siri On-Device Deep Learning-Guided Unit Selection Text-to-Speech System” [9] T. Capes, P. Coles, A. Conkie, L. Golipour, A. Hadjitarkhani, Q. Hu, N. Huddleston, M. Hunt, J. Li, M. Neeracher, K. Prahallad, T. Raitio, R. Rasip…

Because then it wouldn't be an Apple™ iNovation™.

Re: Deep Learning for Siri’s Voice

#29
post #8

It might seem silly, but I'm looking forward to the first AI talk therapist. Most of the benefit of therapy is the talking, so it's not as crazy as it sounds.

At my little company, iCouch, we have experimented with such things, but to actually make it effective — that requires a good amount of capital — capital that is very difficult to raise. I would need to hire 3 full time people just for the AI project and potentially more.

The VC world is interested in “traction” and not novel tech which means we have to divert effort into growing customers for our mental health practice management system to get “traction” before we can spend any notable time building AI therapists. As much as VCs talk about “looking for innovation” they really aren’t. They are just looking at current growth/revenue. The days of building something amazing and monetizing later seem to be over for all except for founders with marquee names.

We could launch AI therapists within a year, but in the meantime, I have to pay my team. So we are forced to subsidize moonshot R&D with our existing sales — but that is hard to do since existing sales have to finance customer acquisition. Finding an additional $500k per year to make AI therapy viable is impossible for us.

We are in a catch 22. The first question from nearly every investor’s mouth: “how many paid users do you have?” Not, what technology do you have or can develop that is truly disruptive. We could start preparing AI therapy tomorrow for a Summer 2018 launch if we could afford it. But if we diverted resources to that, we’d be out of business long before launch. Clinically effective AI therapy isn’t a weekend side project.

Re: Deep Learning for Siri’s Voice

#30
post #27

The iOS 11 Siri sounds like it's a real person talking, it's amazing. Does anyone know if there's an open-source TTS library available with such quality (or if anyone is working on one, from this paper)? I would love to have my home speakers announce things in this voice.

She sounds younger to me, but very natural sounding. Will be interesting to see how Siri on Home Pod works out.

Yes. And the way it answers some questions, it also seems to be going for a more casual, enthusiastic sort of vibe. It wasn't clear to me to what degree the changes apply outside of the female American voice.
Post reply on HN