Live data from Hacker News

Generate audiobooks from E-books with Kokoro-82M

claudio.uk

181–190 of 255 posts

Re: Generate audiobooks from E-books with Kokoro-82M

#181
post #33

Can anyone recommend an open source option that would allow training on a custom voice (my own, so I'd be able to record as many snippets as it needed to train on) to allow me to use it for TTS generation without sharing it off my machine? Edit: I'll wait to see if any recommendations get made here, if not I might give this one a go: https://github.com/coqui-ai/TTS

There is a fork here https://github.com/idiap/coqui-ai-TTS 'coqui-tts'

Though according to the TTS leaderboard, Fish Speech https://github.com/fishaudio/fish-speech and Kokoro are higher.

https://huggingface.co/hexgrad/Kokoro-82M

https://huggingface.co/fishaudio/fish-speech-1.5

Re: Generate audiobooks from E-books with Kokoro-82M

#182
post #79
post #67

Earlier quoted context omitted.

that's the thing. it's not just for accessibility. anything not already narrated is a fair target for TTS. i don't have time to sit down and read books. all reading is done on the go, while getting around or doing daily routines at home. i have a small book that i am reading now, which should take a few hours to finish, but in the time i manage to get done reading it i will probably have listened to two or three audi…

You don't choose to spend your time reading books. You probably roll your eyes when someone tells you they don't have time for some activity you deem valuable. This is the 'no time to exercise' debate in a different shape. They are also different activities, with audio it's easier to listen to more but retention is usually lower. Not casting any elitist "you need to read" bullshit by the way, but find it odd to defin…

This is a weird comment. They are just saying why they prefer audiobooks thus why general TTS is useful for them.

Why are you trying to argue about their preference? They didn't cast any judgement on others with different preferences.

This is nothing like “no time for exercise”.

It's more like "I have no time (preference) to fire up the wood stove so I use microwave" and then you come in with "wow so you roll your eyes at us fire stove users?"

Re: Generate audiobooks from E-books with Kokoro-82M

#185
post #52
post #47

Earlier quoted context omitted.

There was some hope with the rise of equestrianism that people will go back to be able to shoe horses. Guess it was just a matter of time till someone figured out how to use "cars" to resume encouraging being unable to to a basic farrier job.

Except cars were faster than horses, while audio or video content is much slower than reading.

You can multitask with audio content, so you can consume content when you can't sit down to read. And you can even potentially consume more volume like on a long daily commute.

It's not the case that it's worse.

Re: Generate audiobooks from E-books with Kokoro-82M

#187

For anyone looking for an easier alternative (and one without the bugs the author describes, such as skipping some prefaces or failing to detect some chapters), Voice Dream Reader on iOS (and macOS) handles .epub and other e-books just fine and supports a variety of built-in and external voices.

ElevenLabs Reader is the same thing, but much higher quality voices for free. I've lost my place a few times so it's not quite as reliable as VoiceDream. But you aren't paying an expensive subscription with mediocre voices.

Re: Generate audiobooks from E-books with Kokoro-82M

#188

On the one hand, this is very convenient. Probably cool for some non-fiction. On the other, some of my favorite audio books all stood out because the narrator was interpreting the text really well, for example by changing the pacing during chaotic moments. Or those audiobooks with multiple narrators and different voices for each character. Not to mention that sometimes the only cue you get for who's speaking during d…

Would a "better" AI would do a "better" narration with a better understanding of the text? Of course that it would imply a different (and far bigger?) model.

Anyway, even if in theory it might, in practice things may end even worse than doing it with a monotone voice.

Re: Generate audiobooks from E-books with Kokoro-82M

#189

The quality is great (amazing even), but I can't listen to AI generated voices for more than 1 minute. I don't know why, I just don't like it. I immediately skip the video on youtube if the voice is AI generated. Might be because our brains try to 'feel' the speaker, the emotion, the pauses, the invisible smile, etc. No doubt models will improve and will be harder to identify as AI generated, but for now, as with dif…

Among other things, what I don't like is the hallucinated stress. Take the classic example of:

> I never said she stole my money

It can have 7 different meanings based on which word you stress out.

The new AI voices sound very natural at a shallow level, but overall pronounce things in odd ways. Not quite wrong, but subtly unnatural which introduces some cognitive load.

Old TTS systems with their monotonic voices are less confusing, but sound very robotic.

Re: Generate audiobooks from E-books with Kokoro-82M

#190
post #33

Can anyone recommend an open source option that would allow training on a custom voice (my own, so I'd be able to record as many snippets as it needed to train on) to allow me to use it for TTS generation without sharing it off my machine? Edit: I'll wait to see if any recommendations get made here, if not I might give this one a go: https://github.com/coqui-ai/TTS

I wrote this a while ago about xTTSv2 mixed with Nvidia's Nemo. Maybe it kicks off your journey.

https://jdsemrau.substack.com/p/teaching-your-agent-to-speak...

Post reply on HN