Live data from Hacker News

WaveNet launches in the Google Assistant

deepmind.com

1–10 of 122 posts

Re: WaveNet launches in the Google Assistant

#2
I'm interested in using similar generative adversarial networks to reduce artifacting in video streams. For example, highly compressed streams tend to show blocking artifacts on dark scenes, gradients, and static that could be smoothed in the decoder.

I haven't actually done much about it yet, but I'm interested.

Re: WaveNet launches in the Google Assistant

#3
To be honest, it still obviously sounds machine generated. I guess it's a slight improvement, but the examples shown do not include any challenging words or phrases. We've been able to generate adequate sounding speech for simple phrases for quite some time. I bet this was really fun to work on however.

Re: WaveNet launches in the Google Assistant

#4

To be honest, it still obviously sounds machine generated. I guess it's a slight improvement, but the examples shown do not include any challenging words or phrases. We've been able to generate adequate sounding speech for simple phrases for quite some time. I bet this was really fun to work on however.

I thought it's just me. I remember the first WaveNet demo sounded significantly more natural than the non-WaveNet stuff. But now they sounded almost the same. The main difference is that with WaveNet you don't hear that robotic tone at the end of a phrase, that's typical of old TTS technologies. But the way the phrase is spoken still sounds like a machine said it, rather than a human.

I wonder what compromises they made to improve the performance by 1,000x. There have to be some.

Re: WaveNet launches in the Google Assistant

#6

To be honest, it still obviously sounds machine generated. I guess it's a slight improvement, but the examples shown do not include any challenging words or phrases. We've been able to generate adequate sounding speech for simple phrases for quite some time. I bet this was really fun to work on however.

If you ever listen to people try to record sound that's clear and precise, they actually sound fairly robotic. See this Google 20% Project where they explore Google Assistant's voice creation: https://youtu.be/qnGNfz7JiZ8?t=5m23s

WaveNet is probably modeling the source data very well. It sounds like they just need more data with emotion and inflection, rather than having source data that is optimized for monotonicity and precision.

Re: WaveNet launches in the Google Assistant

#7

To be honest, it still obviously sounds machine generated. I guess it's a slight improvement, but the examples shown do not include any challenging words or phrases. We've been able to generate adequate sounding speech for simple phrases for quite some time. I bet this was really fun to work on however.

I guess the hope is that even if it doesn't sound much better, this will open new doors, such as being able to iterate faster on new languages, new voices and new speaking styles (singing, whispering, etc).

EDIT: A big example of that is the Japanese example at the bottom. English has had far more effort put into the old model, so it's already pretty good. But the difference between the old and new Japanese voice is really striking, and they were most likely able to make the new one much faster.

Re: WaveNet launches in the Google Assistant

#8

To be honest, it still obviously sounds machine generated. I guess it's a slight improvement, but the examples shown do not include any challenging words or phrases. We've been able to generate adequate sounding speech for simple phrases for quite some time. I bet this was really fun to work on however.

The Japanese one is miles ahead of the previous one however.
Post reply on HN