Live data from Hacker News

StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

github.com

191–200 of 245 posts

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#191
post #179

Out of curiosity - to folks that have had success with this... This voice cloning is... nothing like XTTSv2, let alone ElevenLabs. It doesn't seem to care about accents at all. It does pretty well with pitch and cadence, and that's about it. I've tried all kinds of different values for alpha, beta, embedding scale, diffusion steps. Anyone else have better luck? Sure it's fast and the sound quality is pretty good, but…

See my previous comment about this point. ElevenLabs are based on Tortoise-TTS which was already pre-trained on millions of hours of data, but this one was only trained on LibriTTS which was 500 hours at best. XTTS was also trained with probably millions of speakers in more than 20 languages. If you have seen millions of voices, there are definitely gonna be some of them that sound like you. It is just a matter of tr…

> It is just a matter of training data, but it is very difficult to have someone collect these large amounts of data and train on it.

It's really not that difficult, they are trained mostly on audiobooks and high quality audio from yt videos. If we talk about EV model then we are talking about around 500k hours of audio, but Tortoise-TTS is only around 50k from what I remember.

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#193

I really want to try this but making the venv to install all the torch dependencies is starting to get old lol. How are other people dealing with this? Is there an easy way to get multiple venvs to share like a common torch venv? I can do this manually but I'm wondering if there's a tool out there that does this.

> is starting to get old lol. If it's starting to get old, then this means that an LLM like Copilot should be able to do it for you, no?

I mean that I already have like 10 different torch venvs for different projects all with various pinned versions and CUDA variants.

Still worth the trade-off of not having to deal with dependency hell, but you start to wonder if there is a better way. All together this is many GBs of duplicated libs, wasted bandwidth and compute.

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#194

Earlier quoted context omitted.

This might be silly because of how few people it benefits, but could it be broken up on to multiple 8GB cards on the same system?

Yes, it absolutely could. You're right that this configuration is rare. Although people have been putting together machines with multiple 24GB cards in order to split and run larger models like llama2-70B.

The latest large models are 120B and 100k context such as Goliath and Tess XL

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#195
post #179

Out of curiosity - to folks that have had success with this... This voice cloning is... nothing like XTTSv2, let alone ElevenLabs. It doesn't seem to care about accents at all. It does pretty well with pitch and cadence, and that's about it. I've tried all kinds of different values for alpha, beta, embedding scale, diffusion steps. Anyone else have better luck? Sure it's fast and the sound quality is pretty good, but…

See my previous comment about this point. ElevenLabs are based on Tortoise-TTS which was already pre-trained on millions of hours of data, but this one was only trained on LibriTTS which was 500 hours at best. XTTS was also trained with probably millions of speakers in more than 20 languages. If you have seen millions of voices, there are definitely gonna be some of them that sound like you. It is just a matter of tr…

What's your basis for the claim that they are based on TorToiSe? I have seen this claim made (and rebutted) many times.

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#196
post #54

I really want to try this but making the venv to install all the torch dependencies is starting to get old lol. How are other people dealing with this? Is there an easy way to get multiple venvs to share like a common torch venv? I can do this manually but I'm wondering if there's a tool out there that does this.

Can relate to this problem a lot. I have considered starting using a Docker dev container and making a base image for shared dependencies which I then can customize in a dockerfile for each new project, not sure if there's a better alternative though.

Yeah there is the official Nvidia container with torch+cuda pre-installed that some projects use.

I feel more projects should start with that as the base instead of pinning on whatever variants. Most aren't using specialized CUDA kernels after all.

Suppose there's the answer, just pick the specific torch+CUDA base that matches the major version of the project you want to run. Then cross your fingers and hope the dependencies mesh :p.

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#197

Earlier quoted context omitted.

Sorry but it was pretty obvious what he meant.

It's not, at all. He could have meant speed, text, audio, words, or phonemes, with least probably images. He probably didn't mean phonemes or he wouldn't be asking. He probably didn't mean arbitrarily slicing 'real' audio and stitching on fake audio - he made repeated references to a video game. He probably didn't mean inpainting and outpainting imagery, even though he made reference to a video game, because its an a…

Inpainting and outpainting of images is when the model generates bits inside or outside the image that don't exist. By analogy he was talking about generating sound inside (I.e. filling gaps) or outside (extrapolating beyond the end) the audio.

I don't know why you would think he was talking about inpainting images, words. This whole discussion is about speech synthesis.

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#198

Earlier quoted context omitted.

What you're not considering here is that a large majority of this industry is made up of no-name voice actors who have a pleasant (but perfectly substitutible) voice which is now something that AI can do perfectly and at a fraction of the price. Sure, celebrities and other well-known figures will have more to gain here as they can license out their voice; but the majority of voice actors won't be able to capitalize o…

No, I am. I work with them, and I've been one (am one, rarely). I listed just one possible use, but I also see voice cloning and advanced TTS expanding access for evocative instruction, as an aid to study style and expand range. Don't be afraid on their behalf. The dooming you're talking about applied to every one of the technological changes I already listed, and we employ more performers and artists today than ever…

TTS actually allows scope for far more different artists' likenesses to be incorporated. An book can be read with all the characters having a different voice entirely. This is difficult currently and relies on the skill of the performer.

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#199
This is really harmful and unethical work. It will be used to hurt millions of elderly people with scams. That's the real application that will happen 100x more than anything else. It's unethical and harmful to release tools that will be overwhelmingly used to hurt elderly people. What they should do about it is: Stop releasing models. Only release a service so that scammers will not use it. Also, only released audio that is watermarked, so that apps can tell that a phone call might be a scam. When they share models with researchers, use previous best practices: post a Google Form to request access.

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#200
post #199

This is really harmful and unethical work. It will be used to hurt millions of elderly people with scams. That's the real application that will happen 100x more than anything else. It's unethical and harmful to release tools that will be overwhelmingly used to hurt elderly people. What they should do about it is: Stop releasing models. Only release a service so that scammers will not use it. Also, only released audio…

Millions of elderly people are already getting scammed by overseas call centers so unless we do something more significant this tech will not make one iota of a difference.
Post reply on HN