Earlier quoted context omitted.
Great stuff, took a look through the README but... what are the minimum hardware requirements to run this? Is this gonna blow up my CPU / harddrive?
Not sure. The only inference demos are colab notebooks. The models are approx 700mb each so I imagine it will run on modest gpu
StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
11–20 of 245 posts
Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#12- synthesize an avatar using stablediffusion
- synthesize conversation with llama
- synthesize the voice with this text thing
soon
- VR
- Video
wild times!
Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#13We're now at "free, local, AI friend that you can have conversations with on consumer hardware" territory. - synthesize an avatar using stablediffusion - synthesize conversation with llama - synthesize the voice with this text thing soon - VR - Video wild times!
Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#14> MIT license > Before using these models, you agree to [...] No, this is not MIT. If you don't like MIT license then feel free to use something else, but you can't pretend this is open source and then attempt to slap on additional restrictions on how the code can be used.
The wording they currently use suggests that this additional license requirement applies not only to their pre-trained models.
Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#15We're now at "free, local, AI friend that you can have conversations with on consumer hardware" territory. - synthesize an avatar using stablediffusion - synthesize conversation with llama - synthesize the voice with this text thing soon - VR - Video wild times!
Which consumer gpu runs llama 70B?
MacBook Pro M3 Max.
Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#16> MIT license > Before using these models, you agree to [...] No, this is not MIT. If you don't like MIT license then feel free to use something else, but you can't pretend this is open source and then attempt to slap on additional restrictions on how the code can be used.
I think you mis-parsed the disclaimer. It's just warning people that cloned voices come with a different set of rights to the software (because the person the voice is a clone of has rights to their voice).
Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#17We're now at "free, local, AI friend that you can have conversations with on consumer hardware" territory. - synthesize an avatar using stablediffusion - synthesize conversation with llama - synthesize the voice with this text thing soon - VR - Video wild times!
Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#18> MIT license > Before using these models, you agree to [...] No, this is not MIT. If you don't like MIT license then feel free to use something else, but you can't pretend this is open source and then attempt to slap on additional restrictions on how the code can be used.
Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech
#19Also, is no one clicking on the audio links? There are some... questionable ones... and I'm pretty sure lots of mistakes.