The prosody and and continuity of the speech is dramatically improved. This is hard to do and very impressive (especially given that it is being done on-device). Personally, I'm less pleased with the actual new voice itself, although that is more a subjective judgment. After listening to many hundreds of voice talent auditions for Alexa, it's hard to step back from that level of pickiness.
As I indicated in another comment, the visual that the voice (together with other tweaks in some of Siri's responses) suggests to me is a perky twenty-something.
I actually tend to generally prefer some of the female British accents in several current TTS systems. (Amy is probably my favorite Polly voice.) Perhaps as an American, the robotic-ness doesn't seem quite as obvious or grating.
I don't like the higher pitch/sharp tone from iOS 11. I like a warmer and deeper tone in iOS 10. I feel like having a more mature/experience assistant.
This just made me realize that every time you see a strong AI in fiction, it still has a computer-sounding voice. If we ever develop strong AI, we will probably already have perfectly natural speech synthesis. And if not, the AI could develop it for us.
But I suppose an AI might choose to use a computer-sounding voice to remind us that it is a computer. Kind of like those inaccurate sound effects in movies - they have become so common that it seems more wrong to omit them. (TV Tropes calls this "The Coconut Effect".)
Good blog post and audio samples notwithstanding, annoying that they don't put the paper on Arxiv. As they themselves point to in the blog post, the learning architecture was introduced in 2014's "Deep mixture density networks for acoustic modeling in statistical parametric speech synthesis" so it's not clear how much of this is just good engineering vs novel research.
The big difference here is the application to unit selection synthesis as opposed to parametric synthesis.
This just made me realize that every time you see a strong AI in fiction, it still has a computer-sounding voice. If we ever develop strong AI, we will probably already have perfectly natural speech synthesis. And if not, the AI could develop it for us. But I suppose an AI might choose to use a computer-sounding voice to remind us that it is a computer. Kind of like those inaccurate sound effects in movies - they hav…
I recommend watching the scifi film "Her", it has a different take on this.
The iOS 11 Siri sounds like it's a real person talking, it's amazing. Does anyone know if there's an open-source TTS library available with such quality (or if anyone is working on one, from this paper)? I would love to have my home speakers announce things in this voice.
I'd love to have my Instapaper articles read to me in that TTS voice.
Hopefully it gets ported to MacOS's say CLI utility. I typically use that with `pbpaste | say` to read my articles.
Their TV service does this too on Google Fiber. It sucks, it's a tire fire. It is laggy, wedges until you have to reboot the network and TV box. It's a horrible experience. I really wish they would stop trying to shove everything into the cloud. Seriously Google, STOP. JUST STOP!
Google's essentially a cloud company, any devices are just remote terminals. I doubt they'll ever change that.
In an increasingly more connected world, it makes sense.
2-3 years ago, Google Docs and Chromebooks were a pain to use for most purposes.
A research paper published by Apple? About Siri?! Unheard of! Last time I was at an NLP conference wth Apple employees they wouldn't say anything about how Siri speech worked, despite being very inquisitive about everyone else's publications. Good to see some change.
It's probably safe to assume a lot of that was due to some/most of Siri being licensed from Nuance initially. I mean, who wants to talk about a new product, which most people think is brand new and entirely innovative, just to say "Oh yeah, we paid someone else to work with us to create it." Not that there's anything wrong with that and it certainly seems like Apple has been investing in-house pretty heavily in recen…
I think it has more to do with the fact that they are finally starting to allow their researchers to publish. They were platinum sponsors at INTERSPEECH 2017 this week, and actually published a paper there. I'm pretty sure that was the first time _ever_ despite their recruiters showing up every year.
Good blog post and audio samples notwithstanding, annoying that they don't put the paper on Arxiv. As they themselves point to in the blog post, the learning architecture was introduced in 2014's "Deep mixture density networks for acoustic modeling in statistical parametric speech synthesis" so it's not clear how much of this is just good engineering vs novel research.
The paper was more than likely embargoed until the talk they gave about it was over. They're introducing some new things that they probably didn't want to release details on before they publicly made a statement.
The prosody and and continuity of the speech is dramatically improved. This is hard to do and very impressive (especially given that it is being done on-device). Personally, I'm less pleased with the actual new voice itself, although that is more a subjective judgment. After listening to many hundreds of voice talent auditions for Alexa, it's hard to step back from that level of pickiness.
As I indicated in another comment, the visual that the voice (together with other tweaks in some of Siri's responses) suggests to me is a perky twenty-something. I actually tend to generally prefer some of the female British accents in several current TTS systems. (Amy is probably my favorite Polly voice.) Perhaps as an American, the robotic-ness doesn't seem quite as obvious or grating.
I also prefer the female British accents but that's exactly what excites me and is so awesome about this. These aren't just samples that are being stitched together anymore. This "learning" that is being done can be applied later to any of the voices in Apple's catalog. Once they get the data of the synthesis out there, they'll more than likely update all the languages and intonations to match. I would imagine that the biggest hurdle with this is that different languages and accents have different nuances. As with most things, they're just starting with English and then will move everything over to all the other options, including the British accents. I don't think we're too far off from a future where you'll be able to pick the age, gender, and voice of your assistant in the same way that characters selection is done in most modern video games.