I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…
ElevenLabs has a much more convincing voice model
VibeVoice: A Frontier Open-Source Text-to-Speech Model
71–80 of 177 posts
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#72I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…
ElevenLabs has a much more convincing voice model
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#73Is there a current, updated list (ideally, a ranking) of the best open weights TTS models? I'm actually more interested in STT (ASR) but the choices there are rather limited.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#74Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#75Ah, yes, the Furious 7 soundtrack. Definitely something everyone recalls
The most popular song of the year from one of the most popular movie franchises that had been in the global news due to the death of its star. Probably the most memorable song from a soundtrack of the century so far.
edit: I had forgotten about Jai Ho (Slumdog Millionaire) and Lose Yourself (8 mile)
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#76I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…
Is there any better model you can point at? I would be interested in having a listen. There are people – and it does not matter what it's about – that will overstate the progress made (and others will understate it, case in point). Neither should put a damper on progress. This is the best I personally have heard so far, but I certainly might have missed something.
However Kokoro-82M is an absolute triumph in the small model space. It curbstomps models 10-20x its size in terms of quality while also being runnable on like, a Raspberry Pi. It’s the kind of thing I’m surprised even exists. Its downside is that it isn’t super expressive, but the af_heart voice is extremely clean, and Kokoro is way more reliable than other TTS models: It doesn’t have the common failure mode where you occasionally have a couple extra syllables thrown in because you picked a bad seed.
If you want something that can do convincing voice acting, either pay for ElevenLabs or keep waiting. If you’re trying to build a local AI assistant, Kokoro is perfect, just use that and check the space again in like 6 months to see if something’s beaten it. https://huggingface.co/hexgrad/Kokoro-82M
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#77Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#78Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#79Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#80Is there a current, updated list (ideally, a ranking) of the best open weights TTS models? I'm actually more interested in STT (ASR) but the choices there are rather limited.
Generally if a model is trending on that page, there’s enough juice for it to be worth a try. There’s a lot of subjective-opinion-having in this space, so beyond “is it trending on HF” the best eval is your own ears. But if something is not trending on HF it is unlikely to be much good.