Out of curiosity - to folks that have had success with this... This voice cloning is... nothing like XTTSv2, let alone ElevenLabs. It doesn't seem to care about accents at all. It does pretty well with pitch and cadence, and that's about it. I've tried all kinds of different values for alpha, beta, embedding scale, diffusion steps. Anyone else have better luck? Sure it's fast and the sound quality is pretty good, but…
See my previous comment about this point. ElevenLabs are based on Tortoise-TTS which was already pre-trained on millions of hours of data, but this one was only trained on LibriTTS which was 500 hours at best. XTTS was also trained with probably millions of speakers in more than 20 languages. If you have seen millions of voices, there are definitely gonna be some of them that sound like you. It is just a matter of tr…
It's really not that difficult, they are trained mostly on audiobooks and high quality audio from yt videos. If we talk about EV model then we are talking about around 500k hours of audio, but Tortoise-TTS is only around 50k from what I remember.