Live data from Hacker News

Google Cloud Text-To-Speech Powered by DeepMind WaveNet Technology

cloudplatform.googleblog.com

91–100 of 122 posts

Re: Google Cloud Text-To-Speech Powered by DeepMind WaveNet Technology

#91

For any Google devs lurking out there, it doesn't seem to work at all in Firefox on Windows. It looks like it has something to do with custom web components with the following message: ReferenceError: customElements is not defined Also apparently some assertion errors with webcomponents (minified so line numbers not useful).

Doesn't work in Safari. Doesn't work in Firefox. Doesn't work in Edge. I don't want to be unreasonable, but Google used to at least generally support the idea of the open web. There are a bunch of different UAs out there; while I accept it's more challenging to support some than others, it doesn't seem unreasonable to expect a product launch from a large-scale web company should at the very least give us an error mes…

> I don't want to be unreasonable, but Google used to at least generally support the idea of the open web.

This is a Machine Learning product, I don't think anyone at Google, at least as part of this team, is trying to "get you" or destroy the open web or something. This isn't even a case of Google using something non-standard -- WebComponents is part of the standard, you can even see it in Mozilla's MDN [0]. Firefox, Safari, Edge, et al simply haven't implemented it yet (or landed in stable). Is that somehow also Chrome's fault?

Filing a bug report is good, but ranting on HN about how this is a sign of Google trying to steal the open Internet is at best unnecessary, and absolutely unreasonable.

Coincidently, I'm working on an application that uses complicated SVG with CSS animations, and I've spent a ton of time optimizing it. I've never tested it outside of Chrome before today. To my surprise, while everything works fast in Chrome, in Safari it's bearable, but in FF it's simply too laggy to use. Now, I probably won't ever get to fix the performance issues in FF and Safari, simply because I don't have the time. Am I also out there trying to destroy the open web? Maybe I'm just bad, not evil.

[0]: https://developer.mozilla.org/en-US/docs/Web/Web_Components

Re: Google Cloud Text-To-Speech Powered by DeepMind WaveNet Technology

#92

Am I wrong in thinking that the cost of generating (realistic-sounding, learned-model) speech on commodity hardware will be near-zero soon, largely negating the value of a SaaS? I've been waiting a long time for decent sounding open source TTS software for narrating books to me, and now with deep learning it's either here or very near here, and the hardware is going to keep getting more performant at the same price.…

"... the hardware is going to keep getting more performant at the same price" That is the hope but there are no guarantees. Perhaps specialized hardware can pick up where Moore's Law has tapered off.

I don't mean general purpose hardware, I mean specifically neural hardware, whether that means GPUs (as we know them), GPUs with special hardware like the "tensor cores", TPUs, "neuromorphic chips" whatever that is, FPGAs becoming part of average computers, or something else.

There's no end in sight to the improvement of neural hardware, not like the wall x86 CPUs have hit anyway.

Re: Google Cloud Text-To-Speech Powered by DeepMind WaveNet Technology

#93
Based on the pricing of $16 per 1 million characters (roughly equal to a 400-500 page book), doesn't this severely threaten the voiceover market place? I just priced the cost of a human voiceover on VoiceBunny.com for a 400-page book and I got an average turnaround time of 90 days / $15K cost vs WaveNet's $16 cost and only 30 mins of computational time. That sounds like an interesting disruptor to me.

Re: Google Cloud Text-To-Speech Powered by DeepMind WaveNet Technology

#94

Based on the pricing of $16 per 1 million characters (roughly equal to a 400-500 page book), doesn't this severely threaten the voiceover market place? I just priced the cost of a human voiceover on VoiceBunny.com for a 400-page book and I got an average turnaround time of 90 days / $15K cost vs WaveNet's $16 cost and only 30 mins of computational time. That sounds like an interesting disruptor to me.

It does if people are willing to listen to the voice for 10 hours+.

I could listen to this voice for a while, but the voice needs more emotion in it before it could be actually useful for long text.

Re: Google Cloud Text-To-Speech Powered by DeepMind WaveNet Technology

#95
post #13

Earlier quoted context omitted.

How so? They don't sell the model itself, they sell the 'tickets' to allow you to take a picture of it. That is not DRM.

Yes, that’s what the OP is saying. Consumers fought against DRM so the business model became “never give them the software and you’ll never need to force DRM on them.”

That's a very cynical interpretation.

Even if Google wants to sell websearch software, there will be so few customers since it would cost billions of dollars to buy machines to run them, you'll need to build DC, get power etc.

And what's wrong with services anyway?

Re: Google Cloud Text-To-Speech Powered by DeepMind WaveNet Technology

#96

Earlier quoted context omitted.

Doesn't work in Safari. Doesn't work in Firefox. Doesn't work in Edge. I don't want to be unreasonable, but Google used to at least generally support the idea of the open web. There are a bunch of different UAs out there; while I accept it's more challenging to support some than others, it doesn't seem unreasonable to expect a product launch from a large-scale web company should at the very least give us an error mes…

> I don't want to be unreasonable, but Google used to at least generally support the idea of the open web. This is a Machine Learning product, I don't think anyone at Google, at least as part of this team, is trying to "get you" or destroy the open web or something. This isn't even a case of Google using something non-standard -- WebComponents is part of the standard, you can even see it in Mozilla's MDN [0]. Firefox…

True, but clearly the marketing page of a feature service for a major cloud vendor should be accessible to as many as possible. There's no real reason something as simple as an audio player to play samples doesn't work well on all browsers.

Re: Google Cloud Text-To-Speech Powered by DeepMind WaveNet Technology

#97
post #31

Does anyone have a GitHub project for epub -> mp3 using this service yet (for automatic audiobook generation)? May make it myself if I have time but curious if anyone already has set it up. EDIT : this is almost exactly their sample application ( https://github.com/GoogleCloudPlatform/python-docs-samples/t... ). Was able to get it working with epubs using pypandoc within the hour. Now just need to make it upload to O…

Shameless plug: https://auditus.cc

Uses Amazon Polly

Re: Google Cloud Text-To-Speech Powered by DeepMind WaveNet Technology

#98
post #94

Based on the pricing of $16 per 1 million characters (roughly equal to a 400-500 page book), doesn't this severely threaten the voiceover market place? I just priced the cost of a human voiceover on VoiceBunny.com for a 400-page book and I got an average turnaround time of 90 days / $15K cost vs WaveNet's $16 cost and only 30 mins of computational time. That sounds like an interesting disruptor to me.

It does if people are willing to listen to the voice for 10 hours+. I could listen to this voice for a while, but the voice needs more emotion in it before it could be actually useful for long text.

Tacotron (also by Google) looks promising in this area https://google.github.io/tacotron/publications/global_style_...

Re: Google Cloud Text-To-Speech Powered by DeepMind WaveNet Technology

#99
post #38

Earlier quoted context omitted.

The machines used to be the thing of value, but now it's the software.

> The machines used to be the thing of value, but now it's the software. Not exactly. The machines were never the value per se, the work they could do was the value. If you can buy the work without buying the machine, it is significantly better.

Yes, but if you could get it for free that would be even better. It's a fantastic business model to charge premium for something that cost almost nothing.

Re: Google Cloud Text-To-Speech Powered by DeepMind WaveNet Technology

#100

Earlier quoted context omitted.

Doesn't work in Safari. Doesn't work in Firefox. Doesn't work in Edge. I don't want to be unreasonable, but Google used to at least generally support the idea of the open web. There are a bunch of different UAs out there; while I accept it's more challenging to support some than others, it doesn't seem unreasonable to expect a product launch from a large-scale web company should at the very least give us an error mes…

> I don't want to be unreasonable, but Google used to at least generally support the idea of the open web. This is a Machine Learning product, I don't think anyone at Google, at least as part of this team, is trying to "get you" or destroy the open web or something. This isn't even a case of Google using something non-standard -- WebComponents is part of the standard, you can even see it in Mozilla's MDN [0]. Firefox…

No, I don’t agree with you.

This is a basic audio playback widget on the web. The web is nominally an open platform. Whichever team built this decided to build it in a way that it would only support Google’s own web browser. It’s not unreasonable to expect a massive web company to build cross-browser support for their user-facing demos, especially in cases where there is obviously no reason that it needs to be incompatible.

Like I said, I appreciate that building cross-browser is not always possible. The difference between you and Google is that you aren’t one of the largest companies in the world, and you don’t publish your own browser.

Post reply on HN