Live data from Hacker News

Pocket TTS: A high quality TTS that gives your CPU a voice

kyutai.org

91–100 of 164 posts

Re: Pocket TTS: A high quality TTS that gives your CPU a voice

#92
post #87

The speed of improvement of tts models reminds me of early days of Stable Diffusion. Can't wait until I can generate audiobooks without infinite pain. If I was an investor I'd short Audible.

Wouldn't audible be perfectly positioned to take advantage of this. They have the perfect setup to integrate this into their offering.

It seems more likely that people will buy a digital copy of the book for a few bucks and then run the TTS themselves on devices they already own.

Re: Pocket TTS: A high quality TTS that gives your CPU a voice

#93

I read this, then realized I needed a browser extension to read my long case study and made a browser interface of this and put this together: https://github.com/lukasmwerner/pocket-reader

You can do the same thing with Firefox' Reader Mode. On Linux you have to set up speech-dispatcher to use your favorite TTS as a backend.Once it is set up, there will be an option to listen the page.

Re: Pocket TTS: A high quality TTS that gives your CPU a voice

#94
post #87

The speed of improvement of tts models reminds me of early days of Stable Diffusion. Can't wait until I can generate audiobooks without infinite pain. If I was an investor I'd short Audible.

I feel like TTS is one of the areas that as evolved the least. Small TTS models have been around for like 5+ years and they've only gotten incrementally better. Giants like ElevenLabs make good sounding TTS but it's not quite human yet and the improvements get less and less each iteration.

Re: Pocket TTS: A high quality TTS that gives your CPU a voice

#95
post #87

The speed of improvement of tts models reminds me of early days of Stable Diffusion. Can't wait until I can generate audiobooks without infinite pain. If I was an investor I'd short Audible.

An all-TTS audiobook offering is just about as appealing as an all-stable-diffusion picture gallery (that is, not at all).

Re: Pocket TTS: A high quality TTS that gives your CPU a voice

#96

I read this, then realized I needed a browser extension to read my long case study and made a browser interface of this and put this together: https://github.com/lukasmwerner/pocket-reader

You can do the same thing with Firefox' Reader Mode. On Linux you have to set up speech-dispatcher to use your favorite TTS as a backend.Once it is set up, there will be an option to listen the page.

Firefox should integrate that in their Reader Mode (the default System Voices are often very un-listable). Would seems like an easy win, and it's a non-AI feature so not polarising.

Re: Pocket TTS: A high quality TTS that gives your CPU a voice

#97
post #87

The speed of improvement of tts models reminds me of early days of Stable Diffusion. Can't wait until I can generate audiobooks without infinite pain. If I was an investor I'd short Audible.

An all-TTS audiobook offering is just about as appealing as an all-stable-diffusion picture gallery (that is, not at all).

Isn’t it more like an art gallery of prints of paintings? The primary art is the text of the book (like the painting in the gallery), TTS (and printing a copy) are just methods of making the art available.

Re: Pocket TTS: A high quality TTS that gives your CPU a voice

#99
post #80

Earlier quoted context omitted.

You can speak one language, switch to another language for one word, and continue speaking in the previous language.

But that's my point. You'll stop, switch, speak, stop, switch, resume. You're not going to be "I was in 東京 yesterday" as a single continuous sentence. It'll have to be broken up to three separate sentences spoken back to back, even for humans.

Not really, most multilinguals switch between languages so seamlessly that you wouldn't even notice it! It even has given birth to new "languages", take for example Hinglish!!

Re: Pocket TTS: A high quality TTS that gives your CPU a voice

#100
post #59

Earlier quoted context omitted.

They didn’t say it was a crazy requirement. They said it was crazy to consider it useless without meeting that requirement.

That doesn't really change what I said though. It isn't crazy to call it useless without some form of ALS either. Given that old school synthesis has been able to do it for like 20 years or so.

How does state of the art matter when talking about usefulness? Is old school synthesis useless?
Post reply on HN