Guys, for those of you who would like to see how Microsoft Cognitive service’s TTS fares when compared to Google’s TTS.
We had launched a Bot which gives voice summary of web content on Messenger, Slack, Telegram & Twitter with In-line audio player on first three. It’s great for sharing audio summary to our visually impaired friends.
I am wondering what their baseline is. They call it "Current Best Non-WaveNet". Quite frankly, Apple's most recent deep learning-based speech synthesis sounds superior, but there aren't enough samples to for a proper comparison: https://machinelearning.apple.com/2017/08/06/siri-voices.htm...
It could just be a matter of opinion, but I prefer both Google's unit selection synthesis, and their WaveNet synthesis. The prosody in Apple's latest method is still annoying, nowhere near as good as the Google models of 2015 and 2016, and not remotely comparable to the WaveNet models. Apple's change in voice talent is an improvement though, and they may have more units than before, which is helpful. I believe their…
It runs on the phone afaik, while speech recognition on the other hand is supported by stuff Google runs in the cloud. Edit: I might be wrong, at the end of a paragraph they say it runs on Google’s TPU cloud infrastructure, though it isn't clear to me whether they just use that for training. Edit 2: I just tried it on my phone. At least stuff like asking it to "Turn on WiFi" works without an internet connection, and…
> I just tried it on my phone. At least stuff like asking it to "Turn on WiFi" works without an internet connection, and yields a TTS response. But this is the status quo. You would not expect Google to disable offline TTS just for slightly improved quality. The real question is, is it running Wavenet offline or the previous version of its TTS engine offline?
Today's announcement is cloud-only. We also support an older algorithm for offline use that's less computationally intensive.
I'm most interested in its potential for audiobooks, especially of the vast array of old and less-common books that don't have human-made audiobook equivalents. I find myself constrained by the limits of audiobook choices, which tends toward best-sellers lists or pop-sci. Current attempts to use text-to-speech to generate audiobooks results in something that's frankly unlistenable. If TTS could get to a "good enough"…
Ancient Kindle Keyboard can text-to-speech most Amazon ebooks with zero hassle (I think authors can disable it). 'Unlistenable' is in the ear of the beholder! My only gripe personally is having a single "narrator" for all content.
I'm most interested in its potential for audiobooks, especially of the vast array of old and less-common books that don't have human-made audiobook equivalents. I find myself constrained by the limits of audiobook choices, which tends toward best-sellers lists or pop-sci. Current attempts to use text-to-speech to generate audiobooks results in something that's frankly unlistenable. If TTS could get to a "good enough"…
Ancient Kindle Keyboard can text-to-speech most Amazon ebooks with zero hassle (I think authors can disable it). 'Unlistenable' is in the ear of the beholder! My only gripe personally is having a single "narrator" for all content.
I agree "unlistenable" is a personal judgement, but there are very rare cases when buying an audiobook, even with a terrible reader, is not a better value/cost ratio than the original Kindle TTS, it was convenient, but not great.
Oh how much I wish Google paid to Morgan Freeman or David Attenborough to have their voice as an option.
Wonder -- in all seriousness -- if the BBC will do Attenborough. If you're British and you've grown up with his documentaries, all other nature voices simply sound wrong.
> I just tried it on my phone. At least stuff like asking it to "Turn on WiFi" works without an internet connection, and yields a TTS response. But this is the status quo. You would not expect Google to disable offline TTS just for slightly improved quality. The real question is, is it running Wavenet offline or the previous version of its TTS engine offline?
Today's announcement is cloud-only. We also support an older algorithm for offline use that's less computationally intensive.
Is the offline one improving also? Google Maps often falls back to it (much more often than necessary for some reason) and it sounds completely different and far worse.