Live data from Hacker News

Generate audiobooks from E-books with Kokoro-82M

claudio.uk

191–200 of 255 posts

Re: Generate audiobooks from E-books with Kokoro-82M

#191

The quality is great (amazing even), but I can't listen to AI generated voices for more than 1 minute. I don't know why, I just don't like it. I immediately skip the video on youtube if the voice is AI generated. Might be because our brains try to 'feel' the speaker, the emotion, the pauses, the invisible smile, etc. No doubt models will improve and will be harder to identify as AI generated, but for now, as with dif…

Among other things, what I don't like is the hallucinated stress. Take the classic example of: > I never said she stole my money It can have 7 different meanings based on which word you stress out. The new AI voices sound very natural at a shallow level, but overall pronounce things in odd ways. Not quite wrong, but subtly unnatural which introduces some cognitive load. Old TTS systems with their monotonic voices are…

erroneous or inappropriate ≠ hallucinated

Re: Generate audiobooks from E-books with Kokoro-82M

#192
post #64

Earlier quoted context omitted.

RIP to future top-enders that would normally have started out on the bottom to middle end.

I'm super opposed to AI, but I see this as a rare positive. As someone already said, the win here is to have a audiobook where one doesn't yet exist. hell, maybe the tables will turn and the scrubs will do the hard work of discovering which titles are popular with an audience, then the ebook industry can capitalize on AI by hiring voice actors to produce proper titles?

Not gonna happen. Once the AI shit is out there, people will have consumed it by the time a real actor can create (and edit) the audiobook.

Re: Generate audiobooks from E-books with Kokoro-82M

#193

The quality is great (amazing even), but I can't listen to AI generated voices for more than 1 minute. I don't know why, I just don't like it. I immediately skip the video on youtube if the voice is AI generated. Might be because our brains try to 'feel' the speaker, the emotion, the pauses, the invisible smile, etc. No doubt models will improve and will be harder to identify as AI generated, but for now, as with dif…

Haven't really been following the latest in TTS ML, but I expected this to be better or at least as good-bad as the stuff you hear on YouTube. Somehow it sounds worse. It really is jarring to listen to any of these ML voices and can't really stand it. Nope out of every video that uses them and can't tell if YouTube never recommends them to me for that reason, or just because the recommendations around what I watch are just so rarely going to be from some low reputation channel.

Take a moment here for a second though and think about it. Even if these voices got to be really good, indistinguishable almost... would I want to listen to it even then? If it was an NPC's generated voice and generated dialogue in a game to help enrich the world building, maybe in that context. On YouTube or with newscasters? Probably not. Audio books? Think I would still rather have it be a real person, because it's like they're reading a story to me and it feels better if it's coming from someone. There's also the unknown factor, where if it's ML generated it's so sterile that the unknowns are kind of gone.

Think about it like this, in the movie industry we had practical effects that were charming in a way. You could think about the physical things that had to occur to make that happen. Movie magic. Now, everything is so CG it's like the magic is gone. Even though you know people put serious hard work into it, there's a kind of inauthenticity and just lack of relevance to the real world that takes something away from it.

It's like a real magician has interesting tricks, while an artificial magician is most likely just a liar.

Still, I grant that it makes some cool things possible and there is potential if things are done right. Some positive mixture of real humans and machine generated stuff so it isn't devoid of anything connected to real life effort.

Re: Generate audiobooks from E-books with Kokoro-82M

#194

On the one hand, this is very convenient. Probably cool for some non-fiction. On the other, some of my favorite audio books all stood out because the narrator was interpreting the text really well, for example by changing the pacing during chaotic moments. Or those audiobooks with multiple narrators and different voices for each character. Not to mention that sometimes the only cue you get for who's speaking during d…

I agree with you, but also want to point out: New authors, self-publishers, can't afford tens of thousands of dollars to get an audiobook recorded professionally... This can limit their distribution. Authors might even choose not to make such version (or lack confidence to record themselves), so AI capable of making a decently passable version would be nice -- something more than reading text blandly. AI in theory co…

You can get narrators to work on a royalty basis.

Re: Generate audiobooks from E-books with Kokoro-82M

#197
post #79

Earlier quoted context omitted.

You don't choose to spend your time reading books. You probably roll your eyes when someone tells you they don't have time for some activity you deem valuable. This is the 'no time to exercise' debate in a different shape. They are also different activities, with audio it's easier to listen to more but retention is usually lower. Not casting any elitist "you need to read" bullshit by the way, but find it odd to defin…

This is a weird comment. They are just saying why they prefer audiobooks thus why general TTS is useful for them. Why are you trying to argue about their preference? They didn't cast any judgement on others with different preferences. This is nothing like “no time for exercise”. It's more like "I have no time (preference) to fire up the wood stove so I use microwave" and then you come in with "wow so you roll your ey…

Two hours before you posted this there was already an admission from me in a sister comment that I came across too judgy and someone else made the point I tried better than myself - not sure how much penitence I need to do but sorry again :)

Re: Generate audiobooks from E-books with Kokoro-82M

#198
post #64

Earlier quoted context omitted.

RIP to future top-enders that would normally have started out on the bottom to middle end.

Virtually every book I want this for has been around for 70+ years and still no high or low quality audiobook has been produced. How long do I have to wait for those aspiring top-enders before an audiobook can be made available?

That has nothing to do with audiobook voice actors and everything to do with copyright and who owns the rights to the book (and whether they believe there's any money to be made selling an audiobook version).

Re: Generate audiobooks from E-books with Kokoro-82M

#200

The quality is great (amazing even), but I can't listen to AI generated voices for more than 1 minute. I don't know why, I just don't like it. I immediately skip the video on youtube if the voice is AI generated. Might be because our brains try to 'feel' the speaker, the emotion, the pauses, the invisible smile, etc. No doubt models will improve and will be harder to identify as AI generated, but for now, as with dif…

> I immediately skip the video on youtube if the voice is AI generated.

I mean, I do that because it's correlated with the content being garbage. If I'm intentionally using it on content I want to consume I expect it to be different, though I haven't gotten around to trying it properly yet so I guess we'll see. (OTOH I already listen to ebooks via pre-AI TTS, so I'm optimistic)

Post reply on HN