Need this for mac
VibeVoice: A Frontier Open-Source Text-to-Speech Model
101–110 of 177 posts
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#102Earlier quoted context omitted.
Is there any better model you can point at? I would be interested in having a listen. There are people – and it does not matter what it's about – that will overstate the progress made (and others will understate it, case in point). Neither should put a damper on progress. This is the best I personally have heard so far, but I certainly might have missed something.
It’s tough to name the best local TTS since they all seem to trade off on quality and features and none of them are as good as ElevenLabs’ closed-source offering. However Kokoro-82M is an absolute triumph in the small model space. It curbstomps models 10-20x its size in terms of quality while also being runnable on like, a Raspberry Pi. It’s the kind of thing I’m surprised even exists. Its downside is that it isn’t s…
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#103I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…
Is there any better model you can point at? I would be interested in having a listen. There are people – and it does not matter what it's about – that will overstate the progress made (and others will understate it, case in point). Neither should put a damper on progress. This is the best I personally have heard so far, but I certainly might have missed something.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#104Earlier quoted context omitted.
I agree. For some reason the female voices are waaay more convincing than the male ones too, which sound barely better than speech synthesis from a decade ago.
Results correlate to investment, and there’s more in synthesizing female coded voices. As for the why female coded voices gets more investments, we all know, only difference is in attitude towards that (the correct answer, of course, is “it sucks”)
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#105they vibecoded their demo website? the text is invisible on Firefox.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#106Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#107does anyone know of recent TTS options that let you specify IPA rather than written words? Azure lets you do this, but something local (and better than existing OS voices) would be great for my project.
My usage is for Chinese, but the phonemes it generated looked very much like IPA.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#108Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#109Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#110What an odd name to me, becaus "Vibe" is, in my mind, equal to somewhat poor quality. Like "Vibe Coding". But that's probably just some bias from my side.