Good quality but unfortunately it is single language English only.
I think they should have added the fact that it's English only in the title at the very least.
91–100 of 164 posts
Good quality but unfortunately it is single language English only.
I think they should have added the fact that it's English only in the title at the very least.
The speed of improvement of tts models reminds me of early days of Stable Diffusion. Can't wait until I can generate audiobooks without infinite pain. If I was an investor I'd short Audible.
Wouldn't audible be perfectly positioned to take advantage of this. They have the perfect setup to integrate this into their offering.
I read this, then realized I needed a browser extension to read my long case study and made a browser interface of this and put this together: https://github.com/lukasmwerner/pocket-reader
The speed of improvement of tts models reminds me of early days of Stable Diffusion. Can't wait until I can generate audiobooks without infinite pain. If I was an investor I'd short Audible.
The speed of improvement of tts models reminds me of early days of Stable Diffusion. Can't wait until I can generate audiobooks without infinite pain. If I was an investor I'd short Audible.
I read this, then realized I needed a browser extension to read my long case study and made a browser interface of this and put this together: https://github.com/lukasmwerner/pocket-reader
You can do the same thing with Firefox' Reader Mode. On Linux you have to set up speech-dispatcher to use your favorite TTS as a backend.Once it is set up, there will be an option to listen the page.
The speed of improvement of tts models reminds me of early days of Stable Diffusion. Can't wait until I can generate audiobooks without infinite pain. If I was an investor I'd short Audible.
An all-TTS audiobook offering is just about as appealing as an all-stable-diffusion picture gallery (that is, not at all).
https://gist.github.com/britannio/481aca8cb81a70e8fd5b7dfa2f...
Earlier quoted context omitted.
You can speak one language, switch to another language for one word, and continue speaking in the previous language.
But that's my point. You'll stop, switch, speak, stop, switch, resume. You're not going to be "I was in 東京 yesterday" as a single continuous sentence. It'll have to be broken up to three separate sentences spoken back to back, even for humans.
Earlier quoted context omitted.
They didn’t say it was a crazy requirement. They said it was crazy to consider it useless without meeting that requirement.
That doesn't really change what I said though. It isn't crazy to call it useless without some form of ALS either. Given that old school synthesis has been able to do it for like 20 years or so.