I made a 100% local voice chatbot using StyleTTS2 and other open source pieces (Whisper and OpenHermes2-Mistral-7B). It responds so much faster than ChatGPT. You can have a real conversation with it instead of the stilted Siri-style interaction you have with other voice assistants. Fun to play with! Anyone who has a Windows gaming PC with a 12 GB Nvidia GPU (tested on 3060 12GB) can install and converse with StyleTTS…
StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
171–180 of 245 posts
Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#172I made a 100% local voice chatbot using StyleTTS2 and other open source pieces (Whisper and OpenHermes2-Mistral-7B). It responds so much faster than ChatGPT. You can have a real conversation with it instead of the stilted Siri-style interaction you have with other voice assistants. Fun to play with! Anyone who has a Windows gaming PC with a 12 GB Nvidia GPU (tested on 3060 12GB) can install and converse with StyleTTS…
Is 12GB the minimum? got an out of memory error with 8GB
Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#173I made a 100% local voice chatbot using StyleTTS2 and other open source pieces (Whisper and OpenHermes2-Mistral-7B). It responds so much faster than ChatGPT. You can have a real conversation with it instead of the stilted Siri-style interaction you have with other voice assistants. Fun to play with! Anyone who has a Windows gaming PC with a 12 GB Nvidia GPU (tested on 3060 12GB) can install and converse with StyleTTS…
Hey modeless. Love it. Is your project open source by any chance? Would love to see it.
Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#174Earlier quoted context omitted.
Short-term could it be configured as push to talk?
Certainly, but then it has little advantage over e.g. ChatGPT voice mode. I guess running locally is an advantage but the voice and answer quality is worse. The much better latency and more natural conversation is what I like about it.
Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#175Earlier quoted context omitted.
Is 12GB the minimum? got an out of memory error with 8GB
Yes, unfortunately these models take a lot of VRAM. It may be possible to do an 8GB version but it will have to compromise on quality of voice recognition and the language model so it might not be a good experience.
Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#176Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#177Earlier quoted context omitted.
Yes, unfortunately these models take a lot of VRAM. It may be possible to do an 8GB version but it will have to compromise on quality of voice recognition and the language model so it might not be a good experience.
This might be silly because of how few people it benefits, but could it be broken up on to multiple 8GB cards on the same system?
Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#178Earlier quoted context omitted.
Certainly, but then it has little advantage over e.g. ChatGPT voice mode. I guess running locally is an advantage but the voice and answer quality is worse. The much better latency and more natural conversation is what I like about it.
Am I wrong to think it would have a couple major advantages? Like using speakers without having to worry about echo cancellation, having a distinct interrupt signal, and still getting all the latency benefits (possibly even more once you get used to it since the conversational style has to assume the end of the user’s sentence instead of knowing the second they let go of the button)
I'd rather use speaker diarization and/or echo cancellation to solve the problem without needing the user to press any buttons.
Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#179Out of curiosity - to folks that have had success with this... This voice cloning is... nothing like XTTSv2, let alone ElevenLabs. It doesn't seem to care about accents at all. It does pretty well with pitch and cadence, and that's about it. I've tried all kinds of different values for alpha, beta, embedding scale, diffusion steps. Anyone else have better luck? Sure it's fast and the sound quality is pretty good, but…
If you have seen millions of voices, there are definitely gonna be some of them that sound like you. It is just a matter of training data, but it is very difficult to have someone collect these large amounts of data and train on it.
Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#180If AI will render some jobs obsolete, I suppose the first one will be audio book narrators and voice actors.
Hardly. Imagine licensing your voice to Amazon so that any customer could stream any book narrated in your likeness without you having to commit the time to record. You could still work as a custom voice artist, all with a "no clone" clause if you chose. You could profit from your performance and craft in a fraction of the time, focusing as your own agent on the management of your assets. Or, you could just keep and…
It's rare that you need a talent like Dan Castellaneta, Mel Blanc, etc.
Secondly, yes, VA licensing will become a thing – but that means that jobs that would previously be available to other lesser known voice actors, because the major players simply didn't have enough time to take those gigs, can no longer take them. A TTSVA can do unlimited recordings.
Thirdly, major studios that would require hundreds of voices for video games and other things don't have to license known voices at all, they can just create generate brand new ones and pay zero licensing fees.