Live data from Hacker News

VibeVoice: A Frontier Open-Source Text-to-Speech Model

microsoft.github.io

151–160 of 177 posts

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#151
VibeVoice-Large is the first local TTS that can produce convincing Finnish speech with little to no accent. I tinkered with it yesterday and was pleasantly surprised at how good the voice cloning is and how it "clones" the emotion in the speech as well.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#152

I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…

The English/Mandarin section was VERY impressive. The accents of both the woman speaking English and the man speaking Chinese were spot on. Both sound very convincingly like they are speaking a second language, which anyone here can hear from the Chinese woman speaking English voice. I'd like to add that the foreigner speaking Chinese was also spot on.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#154

Earlier quoted context omitted.

There's a lot of money and effort spent in satisfying the sexual desires of (predominantly straight) men. There's not typically quite as much interest in doing the same for women. For example I've been looking at models and loras for generating images, and the boards are _full_ of ones that will generate women well or in some particular style. Quite often at least a couple of the preview images for each are hidden be…

I think this is a very lazy kind of cultural analysis. The reason female voices are being chosen over male ones is a little more multifaceted than just SEX. Heterosexual women also tend to prefer female voices over male ones. Female voices are often rated as being clearer, easier to understand, "warmer", etc. Why this is the case is still an open question, but it's definitely more complex than just SEX.

I don't think that this is the only factor, I just suspect that it is _a_ factor.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#156

I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…

Is there any better model you can point at? I would be interested in having a listen. There are people – and it does not matter what it's about – that will overstate the progress made (and others will understate it, case in point). Neither should put a damper on progress. This is the best I personally have heard so far, but I certainly might have missed something.

Elevenlabs v3 (not local)

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#157

I really hope someone within Microsoft is naming their open source coding agent Microsoft VibeCode. Let this be a thing. Its either that or "Lo" then you can have Lo work with Phi, so you can Vibe code with Lo Phi. https://techcommunity.microsoft.com/blog/azure-ai-foundry-bl...

Knowing the history of Microsoft marketing, it will either be called something like "Microsoft Copilot Code Generator for VSCode" or something like "Zunega"...

Microsoft Copilot .NET for Workgroups

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#158
post #76

Earlier quoted context omitted.

It’s tough to name the best local TTS since they all seem to trade off on quality and features and none of them are as good as ElevenLabs’ closed-source offering. However Kokoro-82M is an absolute triumph in the small model space. It curbstomps models 10-20x its size in terms of quality while also being runnable on like, a Raspberry Pi. It’s the kind of thing I’m surprised even exists. Its downside is that it isn’t s…

What is your opinion about F5-TTS or Fish-TTS?

I recently implemented Fish for a project and found it adequate for TTS but wildly impressive in voice cloning. My POC originally required 3-10 audio samples but I removed the minimum because it could usually one shot it.

The model is good, but I will say their inference code leaves a lot to be desired. I had to rewrite large portions of it for simple things like correct chunking and streaming. The advertised expressive keywords are very much hit and miss, and the devs have gone dark unfortunately.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#159
post #145

Earlier quoted context omitted.

I think this is a very lazy kind of cultural analysis. The reason female voices are being chosen over male ones is a little more multifaceted than just SEX. Heterosexual women also tend to prefer female voices over male ones. Female voices are often rated as being clearer, easier to understand, "warmer", etc. Why this is the case is still an open question, but it's definitely more complex than just SEX.

That you consider it sex (rather than gender), is exactly why there’s a preference for female coded voices. Consider where we do hear male recorded voices used as default.

woosh
Post reply on HN