So, when can we expect an open source release of the tech and model?
Never, and it fucking sucks.
WaveNet launches in the Google Assistant
41–50 of 122 posts
Re: WaveNet launches in the Google Assistant
#42I am wondering what their baseline is. They call it "Current Best Non-WaveNet". Quite frankly, Apple's most recent deep learning-based speech synthesis sounds superior, but there aren't enough samples to for a proper comparison: https://machinelearning.apple.com/2017/08/06/siri-voices.htm...
Apple's change in voice talent is an improvement though, and they may have more units than before, which is helpful. I believe their model also works offline, which is a huge plus (though I think Google's prior model works offline as well).
Re: WaveNet launches in the Google Assistant
#43I'm most interested in its potential for audiobooks, especially of the vast array of old and less-common books that don't have human-made audiobook equivalents. I find myself constrained by the limits of audiobook choices, which tends toward best-sellers lists or pop-sci. Current attempts to use text-to-speech to generate audiobooks results in something that's frankly unlistenable. If TTS could get to a "good enough"…
Some of the voices that came out of that were incredible.
Re: WaveNet launches in the Google Assistant
#44So, when can we expect an open source release of the tech and model?
Re: WaveNet launches in the Google Assistant
#45I'm most interested in its potential for audiobooks, especially of the vast array of old and less-common books that don't have human-made audiobook equivalents. I find myself constrained by the limits of audiobook choices, which tends toward best-sellers lists or pop-sci. Current attempts to use text-to-speech to generate audiobooks results in something that's frankly unlistenable. If TTS could get to a "good enough"…
This was kind of the target of the 2016 Blizzard Challenge ( http://festvox.org/blizzard/blizzard2016.html ), as the training data was children's books (clean audio, 'reading voice' prosody, etc). Some of the voices that came out of that were incredible.
Re: WaveNet launches in the Google Assistant
#46I've always wondered why companies don't just take all the close-captioned TV streams and use that as training data for their voice models. Seems like it would create a much more natural sounding voice model (at least as far as humans are accustomed to).
What happens if the CC isn't entirely in sync with the video or audio?
Re: WaveNet launches in the Google Assistant
#47Re: WaveNet launches in the Google Assistant
#48So, when can we expect an open source release of the tech and model?
This is Google, so they'll release a paper and then someone will clone it as an Apache project. Or maybe Linux Foundation these days.
Google is only having any success with this because they are getting training data from the public for free, they should also return the model to the public for free (at least for noncommercial use)
Re: WaveNet launches in the Google Assistant
#49Earlier quoted context omitted.
This is Google, so they'll release a paper and then someone will clone it as an Apache project. Or maybe Linux Foundation these days.
I sure hope so, although I am extremely angered that they don’t release the model. Google is only having any success with this because they are getting training data from the public for free, they should also return the model to the public for free (at least for noncommercial use)
(And I don't think WaveNet is even based on public training data.)
Re: WaveNet launches in the Google Assistant
#50Wow, I hear a huge improvement in the Japanese model: the difference between a robot in person and a young woman on the phone.