Earlier quoted context omitted.
A Mac with a lot of unified RAM can do it, or a dual 3090/4090 setup gets you 48gb of VRAM.
Does this actually work? I had thought that you can't use SLI to increase your net memory for the modal?
StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
81–90 of 245 posts
Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#82Could have interesting prospects for games where you have LLM assuming a character and such TTS giving those NPCs voice.
Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#83Earlier quoted context omitted.
Ah ok, thanks. I tried the other demo.
I tried it. Sounds absolutely nothing like my voice or my wife's voice. I used the same sample files as I used 2 days ago on the Eleven Labs website, and they worked flawlessly there. So this is very, very far from being close to "Eleven Labs quality" when it comes to voice cloning.
Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#84Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#85If AI will render some jobs obsolete, I suppose the first one will be audio book narrators and voice actors.
Just imagine hearing the final novel of ASoIaF narrated by Roy Dotrice and knowing that a royalty went to his family and estate, or if David Attenborough willed the digital likeness of his voice and its performance to the BBC for use in nature documentaries after his death.
The advent of recorded audio didn't put artists out of business, it expanded the industries that relied on them by allowing more of them to work. Film and tape didn't put artists out of business, it expanded the industries that relied on them by allowing more of them to work. Audio digitization and the internet didn't put artists out of business; it expanded the industries that relied on them by allowing more of them to work.
And TTS won't put artists out of business, but it will create yet another new market with another niche that people will have to figure out how to monetize, even though 98% of the revenues will still somehow end up with the distributors.
Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#86The quality is really really INSANE and pretty much unimaginable in early 2000s. Could have interesting prospects for games where you have LLM assuming a character and such TTS giving those NPCs voice.
Currently playing in a golf simulator has a bit of a post-apocalyptian vibe. The birds are cheeping, the grass is rustling, the game play is realistic, but there's not a human to be seen. Just so different from the smacktalking of a real round, or the crowd noise at a big game.
It's begging for some LLM-fuelled banter to be added.
Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#87Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#88If AI will render some jobs obsolete, I suppose the first one will be audio book narrators and voice actors.
Hardly. Imagine licensing your voice to Amazon so that any customer could stream any book narrated in your likeness without you having to commit the time to record. You could still work as a custom voice artist, all with a "no clone" clause if you chose. You could profit from your performance and craft in a fraction of the time, focusing as your own agent on the management of your assets. Or, you could just keep and…
Sure, celebrities and other well-known figures will have more to gain here as they can license out their voice; but the majority of voice actors won't be able to capitalize on this. So this is actually even more perverse because it again creates a system where all assets will accumulate at the top and there won't be any distributions for everyone else.
Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#89If AI will render some jobs obsolete, I suppose the first one will be audio book narrators and voice actors.
Hardly. Imagine licensing your voice to Amazon so that any customer could stream any book narrated in your likeness without you having to commit the time to record. You could still work as a custom voice artist, all with a "no clone" clause if you chose. You could profit from your performance and craft in a fraction of the time, focusing as your own agent on the management of your assets. Or, you could just keep and…
Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#90Earlier quoted context omitted.
It's not called a GAN TTS right? StyleGAN is called what it is because of a "style-based" approach and StyleTTS/2 seems to be doing the same (applying style transfer) through different method (and disentangling style from the rest of the voice synthesis). (Actually, looked at the original StyleTTS paper and it actually even partially uses AdaIN in the decoder, which is the same way that StyleGAN injected style inform…
Yeah no I get this but the naming convention has become so prolific that anyone working in generative space hears "Style " and you should think "GAN". (I work in generative vision btw) My point is not that it is technically right, it is that the name is strongly related with the concept now. Such that if you use a style based network and don't name it StyleX that it's odd and might look like you're trying to claim yo…
(I'm half kidding, I get what you mean, but also, think about it. The alternative is worse.)