Live data from Hacker News

VibeVoice: A Frontier Open-Source Text-to-Speech Model

microsoft.github.io

11–20 of 177 posts

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#12
Very good and I could see how I might believe they are real people if I let my guard down. The male voice sounded a little sedated though and there was a smoothness to it that could be samey over long stretches.

Still not at the astonishing level of Google Notebook text to speech which has been out for a while now. I still can't believe how good that one is.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#13
post #8
post #5

Earlier quoted context omitted.

Whisper is speech-to-text. VibeVoice is text-to-speech.

There is a text-to-speech version of whisper, but IMHO the quality is much worse than the demos of this model.

Are you referring to this?

https://github.com/WhisperSpeech/WhisperSpeech

Or is there some OpenAI official Whisper TTS?

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#14

This is clearly high quality but there's something about the voices, the male voices in particular, which immediately register as computer generated. My audio vocabulary is not rich enough to articulate what it is.

After hearing them myself, I think I know what you mean. The voices get a bit warbly and sound at times like they are very mp3-compressed.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#15

This is clearly high quality but there's something about the voices, the male voices in particular, which immediately register as computer generated. My audio vocabulary is not rich enough to articulate what it is.

I'm no audio engineer either, but those computer voice sound "saw-tooth"y to me.

From what I understand, it's more basic models/techniques that are undersampling, so there is a series of audio pulses which give it that buzzy quality. Better models are produced smoother output.

https://www.perfectcircuit.com/signal/difference-between-wav...

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#17
I really hope someone within Microsoft is naming their open source coding agent Microsoft VibeCode. Let this be a thing. Its either that or "Lo" then you can have Lo work with Phi, so you can Vibe code with Lo Phi.

https://techcommunity.microsoft.com/blog/azure-ai-foundry-bl...

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#18

This is clearly high quality but there's something about the voices, the male voices in particular, which immediately register as computer generated. My audio vocabulary is not rich enough to articulate what it is.

I would describe it as blockly, as if we visualise the sound wave it seems to be without peaks and cut upwards and downwards producing a metallic boxy echo.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#19
post #13
post #8

Earlier quoted context omitted.

There is a text-to-speech version of whisper, but IMHO the quality is much worse than the demos of this model.

Are you referring to this? https://github.com/WhisperSpeech/WhisperSpeech Or is there some OpenAI official Whisper TTS?

Yep, nothing official that I know, but that one is fairly popular so maybe they were referring to it (although AFAIK it's not frontier?)

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#20

Wow. I admit that I am not a native speaker, but this looks (or rather, sounds) VERY impressive and I could mistake it for hearing two people talking.

The giveaway is they will never talk over each other. Only one speaker at a time, consistently.
Post reply on HN