Live data from Hacker News

VibeVoice: A Frontier Open-Source Text-to-Speech Model

microsoft.github.io

161–170 of 177 posts

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#161
Any insight on my the code and the large model were removed? Some copies are floating around and are MIT licensed. In cases like this I do not know why the projects are yanked. If the project was mistakenly released under MIT, copied elsewhere, is any damage control possible by yanking the copies you have control over? Mostly seems like bad PR, if minor.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#162

Any insight on my the code and the large model were removed? Some copies are floating around and are MIT licensed. In cases like this I do not know why the projects are yanked. If the project was mistakenly released under MIT, copied elsewhere, is any damage control possible by yanking the copies you have control over? Mostly seems like bad PR, if minor.

Wondering this too.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#163
post #145

Earlier quoted context omitted.

I think this is a very lazy kind of cultural analysis. The reason female voices are being chosen over male ones is a little more multifaceted than just SEX. Heterosexual women also tend to prefer female voices over male ones. Female voices are often rated as being clearer, easier to understand, "warmer", etc. Why this is the case is still an open question, but it's definitely more complex than just SEX.

That you consider it sex (rather than gender), is exactly why there’s a preference for female coded voices. Consider where we do hear male recorded voices used as default.

How the hell would you determine someone's self assigned social gender based on there voice which is a result of there physical sex.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#164

Is there a current, updated list (ideally, a ranking) of the best open weights TTS models? I'm actually more interested in STT (ASR) but the choices there are rather limited.

Best TTS: VibeVoice, Chatterbox, Dia, Higgs, F5 TTS, Kokoro, Cosy Voice, XTTS-2.

Unmute.sh (same team as Kokoro) gets slept on, but it's really good.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#165

I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…

Is there any better model you can point at? I would be interested in having a listen. There are people – and it does not matter what it's about – that will overstate the progress made (and others will understate it, case in point). Neither should put a damper on progress. This is the best I personally have heard so far, but I certainly might have missed something.

Higgs Audio v2 is currently SOTA in OSS TSS.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#170

Any insight on my the code and the large model were removed? Some copies are floating around and are MIT licensed. In cases like this I do not know why the projects are yanked. If the project was mistakenly released under MIT, copied elsewhere, is any damage control possible by yanking the copies you have control over? Mostly seems like bad PR, if minor.

Ok anyone have a link to the code and weights?
Post reply on HN