I am wondering what their baseline is. They call it "Current Best Non-WaveNet". Quite frankly, Apple's most recent deep learning-based speech synthesis sounds superior, but there aren't enough samples to for a proper comparison: https://machinelearning.apple.com/2017/08/06/siri-voices.htm...
WaveNet launches in the Google Assistant
21–30 of 122 posts
Re: WaveNet launches in the Google Assistant
#22Re: WaveNet launches in the Google Assistant
#23While the 100x speedup sounds impressive, a raw speed number without details regarding hardware is pretty meaningless. I'm guessing they got the 20x realtime speed from running it on their new TPU hardware, which they say can do 180 Tflops. That means you would need 9 Tflops of computing power to run this in realtime -- still pretty far away from running on a phone, or PC for that matter.
In fact that's the title of the article.
Re: WaveNet launches in the Google Assistant
#24To be honest, it still obviously sounds machine generated. I guess it's a slight improvement, but the examples shown do not include any challenging words or phrases. We've been able to generate adequate sounding speech for simple phrases for quite some time. I bet this was really fun to work on however.
If you ever listen to people try to record sound that's clear and precise, they actually sound fairly robotic. See this Google 20% Project where they explore Google Assistant's voice creation: https://youtu.be/qnGNfz7JiZ8?t=5m23s WaveNet is probably modeling the source data very well. It sounds like they just need more data with emotion and inflection, rather than having source data that is optimized for monotonicity…
Re: WaveNet launches in the Google Assistant
#25To be honest, it still obviously sounds machine generated. I guess it's a slight improvement, but the examples shown do not include any challenging words or phrases. We've been able to generate adequate sounding speech for simple phrases for quite some time. I bet this was really fun to work on however.
If you ever listen to people try to record sound that's clear and precise, they actually sound fairly robotic. See this Google 20% Project where they explore Google Assistant's voice creation: https://youtu.be/qnGNfz7JiZ8?t=5m23s WaveNet is probably modeling the source data very well. It sounds like they just need more data with emotion and inflection, rather than having source data that is optimized for monotonicity…
If you work in radio or voiceover you learn really quickly that the voice is so much more complex than people give it credit for. Subtle changes in delivery, timing, inflection, syllabic emphasis, pauses etc... make a massive difference.
Anyone can "talk like a robot" but speaking naturally is way more dynamical than just making the sounds of the words transition smoothly.
I'm not sure if it's way easier or way harder than we're doing it now.
Re: WaveNet launches in the Google Assistant
#26While the 100x speedup sounds impressive, a raw speed number without details regarding hardware is pretty meaningless. I'm guessing they got the 20x realtime speed from running it on their new TPU hardware, which they say can do 180 Tflops. That means you would need 9 Tflops of computing power to run this in realtime -- still pretty far away from running on a phone, or PC for that matter.
They literally said it's launching in the Google Assistant (i.e. running on phones) now. In fact that's the title of the article.
Re: WaveNet launches in the Google Assistant
#27Re: WaveNet launches in the Google Assistant
#28While the 100x speedup sounds impressive, a raw speed number without details regarding hardware is pretty meaningless. I'm guessing they got the 20x realtime speed from running it on their new TPU hardware, which they say can do 180 Tflops. That means you would need 9 Tflops of computing power to run this in realtime -- still pretty far away from running on a phone, or PC for that matter.
They literally said it's launching in the Google Assistant (i.e. running on phones) now. In fact that's the title of the article.