Live data from Hacker News

VibeVoice: A Frontier Open-Source Text-to-Speech Model

microsoft.github.io

61–70 of 177 posts

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#61

I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…

ElevenLabs has a much more convincing voice model

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#62

Unfortunately it's not usable if you're GPU-poor. Couldn't figure out how to run this with an old 1080. I tried VibeVoice-1.5B on my old CPU with torch.float32 and it took 832 seconds to generate a 66 second audio clip. Switching from torch.bfloat16 also introduced some weird sound artifacts in the audio output. If you're GPU-poor the best TTS model I've tried so far is Kokoro. Someone else mentioned in this thread t…

This is ludicrous. macOS has had text-to-speech for ages with acceptable quality, and they never needed energy- and compute-expensive models for it. And it reacts instantly, not after ridiculous delays. I cannot believe this hype about “AI”, it’s just too absurd.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#63
post #42

Looking forward to the day when tts and speech recognition will work on Croatian, or other less prevalent languages. It seems that it's only variants of English, Spanish and Chinese which are somewhat working.

Have you tried Soniox for speech recognition? It supports Croatian. Or are you just looking for self-hosted open-source models? Soniox is very cheap ($0.1/h for async, $0.12/h for real-time) and you get $200 free credits on signup.

https://soniox.com/

Disclaimer: I used to work for Soniox

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#64
post #46

Is there a current, updated list (ideally, a ranking) of the best open weights TTS models? I'm actually more interested in STT (ASR) but the choices there are rather limited.

Click leaderboard in the hamburger menu: https://huggingface.co/spaces/TTS-AGI/TTS-Arena-V2

Is there a way to filter out hosted models? The top three winners currently are all proprietary as far as I can tell.

edit: Ah, there's a lock icon next to the name of each proprietary model.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#66

Unfortunately it's not usable if you're GPU-poor. Couldn't figure out how to run this with an old 1080. I tried VibeVoice-1.5B on my old CPU with torch.float32 and it took 832 seconds to generate a 66 second audio clip. Switching from torch.bfloat16 also introduced some weird sound artifacts in the audio output. If you're GPU-poor the best TTS model I've tried so far is Kokoro. Someone else mentioned in this thread t…

This is ludicrous. macOS has had text-to-speech for ages with acceptable quality, and they never needed energy- and compute-expensive models for it. And it reacts instantly, not after ridiculous delays. I cannot believe this hype about “AI”, it’s just too absurd.

> with acceptable quality

Compared to IBMs Steven Hawking's chair, maybe. But apple tts is not acceptable quality in any modern understanding of SotA, IMO.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#67
post #54

I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…

This is close to SOTA emotional performance, at least the female voices. I trust the human scores in the paper. At least my ear aligns with that figure. With stuff like this coming out in the open, I wonder if ElevenLabs will maintain its huge ARR lead in the field. I really don't see how they can continue to maintain a lead when their offering is getting trounced by open models.

Hmmmm… what is your opinion on the examples showcased here vs the ones on the Dia demo page?

https://yummy-fir-7a4.notion.site/dia

I am not sure why but I find the pacing of the parakeet based models (like Dia) to be much more realistic.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#68

I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…

Is there any better model you can point at? I would be interested in having a listen.

There are people – and it does not matter what it's about – that will overstate the progress made (and others will understate it, case in point). Neither should put a damper on progress. This is the best I personally have heard so far, but I certainly might have missed something.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#69

I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…

It's good but not the best free model. I find Chatterbox to be more realistic with no robot-sounding and better (though not perfect) intonation.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#70
post #46

Is there a current, updated list (ideally, a ranking) of the best open weights TTS models? I'm actually more interested in STT (ASR) but the choices there are rather limited.

Click leaderboard in the hamburger menu: https://huggingface.co/spaces/TTS-AGI/TTS-Arena-V2

That's a highly incomplete comparison
Post reply on HN