Live data from Hacker News

WaveNet launches in the Google Assistant

deepmind.com

81–90 of 122 posts

Re: WaveNet launches in the Google Assistant

#81
post #74
post #59

Earlier quoted context omitted.

Oh man, I can already see the court cases of a certain robot's voice sounding a little too similar to a deceased human's that it should make royalty payments.

Heh, in my original post I wanted to follow up that paragraph with a question: Is someone's voice their IP? Would e.g. Don LaFontaine's family go all J.R.R. Tolkien and forbid the use of his voice in certain contexts? Then again, for now, it seems that it's easy to get away with synthesizing voices of popular characters as long as you don't use copyrighted names: https://acapela-box.com/AcaBox/index.php (choose Engli…

I don't think voices are copyrightable, but the recordings used to analyze them are, and so the resulting voice might be considered a derivative work.

Re: WaveNet launches in the Google Assistant

#82

I am wondering what their baseline is. They call it "Current Best Non-WaveNet". Quite frankly, Apple's most recent deep learning-based speech synthesis sounds superior, but there aren't enough samples to for a proper comparison: https://machinelearning.apple.com/2017/08/06/siri-voices.htm...

I think I read "commercial" somewhere in there. So it'd be "the best you can buy", though not necesarily "the best other competing companies use" (ie: Apple).

Still, they picked one that makes theirs look vastly superior.

Re: WaveNet launches in the Google Assistant

#83
post #74
post #59

Earlier quoted context omitted.

Oh man, I can already see the court cases of a certain robot's voice sounding a little too similar to a deceased human's that it should make royalty payments.

Heh, in my original post I wanted to follow up that paragraph with a question: Is someone's voice their IP? Would e.g. Don LaFontaine's family go all J.R.R. Tolkien and forbid the use of his voice in certain contexts? Then again, for now, it seems that it's easy to get away with synthesizing voices of popular characters as long as you don't use copyrighted names: https://acapela-box.com/AcaBox/index.php (choose Engli…

I don't see it as being very different than using someone's face to promote something.

For example, should a lifelike digital model (or even a still image) of an actor's face still generate royalties for the actor's family after their death? Both cases are using a unique attribute of the person to promote something.

Re: WaveNet launches in the Google Assistant

#84
post #51

Earlier quoted context omitted.

If your co-worker could share how he did this, it would be appreciated (specifically, what apps/code needs to be run).

Not the coworker, but I've been doing this for a couple years. If you're on iOS, the simplest option is Voice Dream Reader [1]. It can read .epub files directly, along with text files and webpages, and integrates with Dropbox, Pocket, Gutenberg, etc. On Android, you can get Voice Dream Reader (though it has less features than on iOS) or @Voice Aloud Reader which is free [2]. [1] http://www.voicedream.com/reader/ [2]…

It also works reasonably well for academic papers as PDF. Equations get butchered, of course, but overall it's not bad (you can set PDF margins so that headers/footers are not read out aloud on every page). I use it to proof-read my own writing, it helps you spot things that spelling/grammer checkers miss.

Re: WaveNet launches in the Google Assistant

#86
post #74
post #59

Earlier quoted context omitted.

Oh man, I can already see the court cases of a certain robot's voice sounding a little too similar to a deceased human's that it should make royalty payments.

Heh, in my original post I wanted to follow up that paragraph with a question: Is someone's voice their IP? Would e.g. Don LaFontaine's family go all J.R.R. Tolkien and forbid the use of his voice in certain contexts? Then again, for now, it seems that it's easy to get away with synthesizing voices of popular characters as long as you don't use copyrighted names: https://acapela-box.com/AcaBox/index.php (choose Engli…

I can't tell if that's Yoda or Cookie Monster.

Re: WaveNet launches in the Google Assistant

#87

I would love something like this for computer notifications, or any sort of automated system notification. For a concrete example, my RC controller has the ability to do voice prompts (e.g. "landing gear down"), and it would be great if I could get a high-quality voice to speak all this. Hell, I'd settle for an API where I could send text and get high-quality voice back. Maybe I can somehow hack the Google Assistant…

If the texts are fixed, why not just pay someone to record them? You can get 100 words for $5.

Mainly because that price is over the "can be bothered" threshold. I'm not going to find someone on fiverr and contract them to record a few words so I can have aural notifications for finished downloads, for example, but I might write a simple script to take the text and store an MP3 with the audio if it tacks a cent on my AWS bill.

Re: WaveNet launches in the Google Assistant

#88
post #72

Earlier quoted context omitted.

One of my co-workers told me in 2012 that he was doing exactly this. He used an ebook reader to download free ebooks from Gutenberg and then the IVONA text to speech engine for Android to listen to them on his drives. He had already finished a few classics like Treasure Island this way. I'm sure things have improved significantly since 2012 so what you're looking for is probably easily done.

Google's "cloud-based" TTS (not the local one, which was terrible) on the Play Books app was pretty good even years ago. However, it was unusable because it would stop as soon as the screen would turn-off. Other ebook readers' read-aloud features worked with the screen off. So thanks for nothing Google!

That's not my experience, as I can turn off the screen and still read the book. I tested it just now.

Re: WaveNet launches in the Google Assistant

#89

I'm most interested in its potential for audiobooks, especially of the vast array of old and less-common books that don't have human-made audiobook equivalents. I find myself constrained by the limits of audiobook choices, which tends toward best-sellers lists or pop-sci. Current attempts to use text-to-speech to generate audiobooks results in something that's frankly unlistenable. If TTS could get to a "good enough"…

You should check out https://getspeechify.com/ - I use it all the time for exactly this. It also has a great team behind it

Re: WaveNet launches in the Google Assistant

#90

I'm most interested in its potential for audiobooks, especially of the vast array of old and less-common books that don't have human-made audiobook equivalents. I find myself constrained by the limits of audiobook choices, which tends toward best-sellers lists or pop-sci. Current attempts to use text-to-speech to generate audiobooks results in something that's frankly unlistenable. If TTS could get to a "good enough"…

I don't have time at the moment to search it out, but in a previous speech-synth thread there's an in depth discussion of this. In short, people were pessimistic on wavenet as audiobook because audiobooks are all about knowing the story and speaking accordingly

But it will only get better and if humans can learn it...
Post reply on HN