Live data from Hacker News

WaveNet launches in the Google Assistant

deepmind.com

31–40 of 122 posts

Re: WaveNet launches in the Google Assistant

#31
post #28

Earlier quoted context omitted.

They literally said it's launching in the Google Assistant (i.e. running on phones) now. In fact that's the title of the article.

Does TTS run on the phone or the cloud?

It runs on the phone afaik, while speech recognition on the other hand is supported by stuff Google runs in the cloud.

Edit: I might be wrong, at the end of a paragraph they say it runs on Google’s TPU cloud infrastructure, though it isn't clear to me whether they just use that for training.

Edit 2: I just tried it on my phone. At least stuff like asking it to "Turn on WiFi" works without an internet connection, and yields a TTS response.

Re: WaveNet launches in the Google Assistant

#33

I am wondering what their baseline is. They call it "Current Best Non-WaveNet". Quite frankly, Apple's most recent deep learning-based speech synthesis sounds superior, but there aren't enough samples to for a proper comparison: https://machinelearning.apple.com/2017/08/06/siri-voices.htm...

I think the voice for the samples in your link still has the problems they talk about in that article.

There are noticeable blips in the speech that sound unnatural, particularly when certain sound combinations are used.

The very first sample with "Bruce Frederick" is clearly off. The intonation and timing between the end of bruce and the beginning of frederick is... mechanical.

There's a similar problem in the OPs link with the non-wavenet English voice 1 when it says "Wavenet".

Those issues are much less apparent in the wavenet voices. Timing problems are less noticeable, intonation problems are less noticeable.

Frankly, the voices there sound VERY good, compared to anything I've heard.

That said, I completely agree that there's not enough samples there to make any real judgement.

Re: WaveNet launches in the Google Assistant

#34
post #28

Earlier quoted context omitted.

Does TTS run on the phone or the cloud?

It runs on the phone afaik, while speech recognition on the other hand is supported by stuff Google runs in the cloud. Edit: I might be wrong, at the end of a paragraph they say it runs on Google’s TPU cloud infrastructure, though it isn't clear to me whether they just use that for training. Edit 2: I just tried it on my phone. At least stuff like asking it to "Turn on WiFi" works without an internet connection, and…

[deleted]

Re: WaveNet launches in the Google Assistant

#36

I'm most interested in its potential for audiobooks, especially of the vast array of old and less-common books that don't have human-made audiobook equivalents. I find myself constrained by the limits of audiobook choices, which tends toward best-sellers lists or pop-sci. Current attempts to use text-to-speech to generate audiobooks results in something that's frankly unlistenable. If TTS could get to a "good enough"…

One of my co-workers told me in 2012 that he was doing exactly this. He used an ebook reader to download free ebooks from Gutenberg and then the IVONA text to speech engine for Android to listen to them on his drives. He had already finished a few classics like Treasure Island this way. I'm sure things have improved significantly since 2012 so what you're looking for is probably easily done.

If your co-worker could share how he did this, it would be appreciated (specifically, what apps/code needs to be run).

Re: WaveNet launches in the Google Assistant

#37

I've always wondered why companies don't just take all the close-captioned TV streams and use that as training data for their voice models. Seems like it would create a much more natural sounding voice model (at least as far as humans are accustomed to).

What happens if the CC isn't entirely in sync with the video or audio?

Re: WaveNet launches in the Google Assistant

#38

Wow, I hear a huge improvement in the Japanese model: the difference between a robot in person and a young woman on the phone.

this was what I noticed too, I don't really have an ear for Japanese, but it sounded as natural as when I have heard it spoken by others. Huge huge improvement.
Post reply on HN