Earlier quoted context omitted.
macOS already has some great intrinsic TTS capability as the OS seems to include a naturally sounding voice. I recently built a similar tool to just run the "say" command as a background process. Had to wrap it in a Deno server. It works, but with Tahoe it's difficult to consistently configure using that one natural voice, and not the subpar voices downloadable in the settings. The good voice seems to be hidden someh…
> The good voice seems to be hidden somehow. How am I supposed to enable this?
Pocket TTS: A high quality TTS that gives your CPU a voice
111–120 of 164 posts
Re: Pocket TTS: A high quality TTS that gives your CPU a voice
#112Earlier quoted context omitted.
Wouldn't audible be perfectly positioned to take advantage of this. They have the perfect setup to integrate this into their offering.
It seems more likely that people will buy a digital copy of the book for a few bucks and then run the TTS themselves on devices they already own.
Re: Pocket TTS: A high quality TTS that gives your CPU a voice
#113Earlier quoted context omitted.
I echo this. For a TTS system to be in any way useful outside the tiny population of the world that speaks exclusively English, it must be multilingual and dynamically switch between languages pretty much per word. Cool tech demo though!
This is a great illustration that nothing you ever do will be good enough without people whining.
Re: Pocket TTS: A high quality TTS that gives your CPU a voice
#114Re: Pocket TTS: A high quality TTS that gives your CPU a voice
#115I'm psyched to see so much interest in my post about Kyutai's latest model! I'm working on part of a related team in Paris that's building off Kutai's research to provide enterprise-grade voice solutions. If anyone building in this space I'd love to chat and share some our upcoming models and capabilities that I am told are SOTA. Please don't hesitate to ping me via the address in my profile.
[1] https://data.norge.no/en/datasets/220ef03e-70e1-3465-a4af-ed...
Re: Pocket TTS: A high quality TTS that gives your CPU a voice
#116Re: Pocket TTS: A high quality TTS that gives your CPU a voice
#117Question: does anyone recommend a TTS that automatically recognizes emotion from the text it self?
"so and so," he
and the verb is not just "said", but "chuckled", or "whispered", or "said shakily", the output is modified accordingly, or if there's an indication that it's a woman speaking it may pitch up during the quotation. It also tries to guess emotive content from textual content, such if a passage reads angry it may try to make it sound angry. That's more hit-and-miss, but when it hits, it hits really well. A very common failure case is, imagine someone is trying to psych themselves up and they say internally "come on, Steve, stand up and keep going", it'll read it in a deeper voice like it was being spoken by a WW2 sergeant to a soldier.
Re: Pocket TTS: A high quality TTS that gives your CPU a voice
#118Earlier quoted context omitted.
It seems more likely that people will buy a digital copy of the book for a few bucks and then run the TTS themselves on devices they already own.
eBooks are much more expensive then an Audible subscription though.
Re: Pocket TTS: A high quality TTS that gives your CPU a voice
#119Re: Pocket TTS: A high quality TTS that gives your CPU a voice
#120Earlier quoted context omitted.
Isn’t it more like an art gallery of prints of paintings? The primary art is the text of the book (like the painting in the gallery), TTS (and printing a copy) are just methods of making the art available.
I think it can be argued that audiobook's add to the art by adding tone and inflection by the reader. To me, what you're saying is the same as saying the art of a movie is in the script, the video is just the method of making it available. And I don't think that's a valid take
So yes, I mostly agree with GP. An audiobook is a different rendering of the same subject. The content is in the text, regardless of whether it's delivered in written or oral form.