Live data from Hacker News

VibeVoice: A Frontier Open-Source Text-to-Speech Model

microsoft.github.io

31–40 of 177 posts

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#31

Earlier quoted context omitted.

Knowing the history of Microsoft marketing, it will either be called something like "Microsoft Copilot Code Generator for VSCode" or something like "Zunega"...

Well don't forget "Microsoft SQL" ;) They'll name something as though they invented it and now have the worse possible way to google it.

For me it doesn't sounds like they invented it but that it's Microsoft version of SQL idk but I hate Microsoft version of anything

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#32

I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…

One of the things this model is actually quite good at is voice cloning. Drop a recorded sample of your voice into the voices folder, and it just works.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#33
post #20

Wow. I admit that I am not a native speaker, but this looks (or rather, sounds) VERY impressive and I could mistake it for hearing two people talking.

The giveaway is they will never talk over each other. Only one speaker at a time, consistently.

Also the lack of stutter and perfect flow of speech are a dead giveaway

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#34
I tried some TTS models a while ago, but I noticed that none of them allowed to put markup statements in the text. For example, it would be nice to do something like:

     Hey look! [enthusiastic] Should we tell the others? Maybe not ... [giggles]
etc.

In fact, I think this kind of thing is absolutely necessary if you want to use this to replace a voice actor.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#35

Earlier quoted context omitted.

Well don't forget "Microsoft SQL" ;) They'll name something as though they invented it and now have the worse possible way to google it.

“Microsoft Word” haha reminds me of the old joke : “Microsoft Works” is an oxymoron.

Oh my goodness, I forgot about "Microsoft Works" you just shot me back in time to the 2000s

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#36

I really hope someone within Microsoft is naming their open source coding agent Microsoft VibeCode. Let this be a thing. Its either that or "Lo" then you can have Lo work with Phi, so you can Vibe code with Lo Phi. https://techcommunity.microsoft.com/blog/azure-ai-foundry-bl...

Knowing the history of Microsoft marketing, it will either be called something like "Microsoft Copilot Code Generator for VSCode" or something like "Zunega"...

GitHub Dotnet Copilot Code Generator for VSC (new)

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#37

This is clearly high quality but there's something about the voices, the male voices in particular, which immediately register as computer generated. My audio vocabulary is not rich enough to articulate what it is.

I would describe it as blockly, as if we visualise the sound wave it seems to be without peaks and cut upwards and downwards producing a metallic boxy echo.

Yeah it sounds super low bitrate to me, reminds me of someone on Bluetooth microphone

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#39

I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…

> (1) I'm not familiar with any Chinese language, so I couldn't really judge the pronunciation of that

https://en.wikipedia.org/wiki/Gell-Mann_amnesia_effect

Post reply on HN