Live data from Hacker News

StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

github.com

11–20 of 245 posts

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#11

Earlier quoted context omitted.

Great stuff, took a look through the README but... what are the minimum hardware requirements to run this? Is this gonna blow up my CPU / harddrive?

Not sure. The only inference demos are colab notebooks. The models are approx 700mb each so I imagine it will run on modest gpu

Would it run in a cheap non-GPU server?

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#13
post #12

We're now at "free, local, AI friend that you can have conversations with on consumer hardware" territory. - synthesize an avatar using stablediffusion - synthesize conversation with llama - synthesize the voice with this text thing soon - VR - Video wild times!

Which consumer gpu runs llama 70B?

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#14
post #4

> MIT license > Before using these models, you agree to [...] No, this is not MIT. If you don't like MIT license then feel free to use something else, but you can't pretend this is open source and then attempt to slap on additional restrictions on how the code can be used.

This bothered me as well. I opened an issue on the repo asking them to consider updating the license file to reflect these additional requirements.

The wording they currently use suggests that this additional license requirement applies not only to their pre-trained models.

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#15
post #13
post #12

We're now at "free, local, AI friend that you can have conversations with on consumer hardware" territory. - synthesize an avatar using stablediffusion - synthesize conversation with llama - synthesize the voice with this text thing soon - VR - Video wild times!

Which consumer gpu runs llama 70B?

Prosumer gear.

MacBook Pro M3 Max.

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#16
post #8
post #4

> MIT license > Before using these models, you agree to [...] No, this is not MIT. If you don't like MIT license then feel free to use something else, but you can't pretend this is open source and then attempt to slap on additional restrictions on how the code can be used.

I think you mis-parsed the disclaimer. It's just warning people that cloned voices come with a different set of rights to the software (because the person the voice is a clone of has rights to their voice).

(Don’t let’s derail the conversation, please, but “disclaimer” is completely the wrong word here. This is a condition of use. A disclaimer is “this isn’t mine” or “I’m not responsible for this”. Disclaimers and disclosures are quite different things and commonly confused, but this isn’t even either of them.)

Re: StyleTTS2 – open-source Eleven-Labs-quality Text To Speech

#19
Why name it Style if it isn't a StyleGAN? Looks like the first one wasn't either. Interesting to see moves away from flows, especially when none of the flows were modern.

Also, is no one clicking on the audio links? There are some... questionable ones... and I'm pretty sure lots of mistakes.

Post reply on HN