Live data from Hacker News

VibeVoice: A Frontier Open-Source Text-to-Speech Model

microsoft.github.io

81–90 of 177 posts

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#82

I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…

Their comments about the singing and background music are odd. It’s been a while since I’ve done academic research, but something about those comments gave me a strong “we couldn’t figure out how to make background music go away in time for our paper submission, so we’re calling it a feature” vibe as opposed to a “we genuinely like this and think its a differentiator” vibe.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#83
post #75

Earlier quoted context omitted.

The most popular song of the year from one of the most popular movie franchises that had been in the global news due to the death of its star. Probably the most memorable song from a soundtrack of the century so far.

I'm Just Ken (Barbie), Skyfall, Let it Go (Frozen), Remember Me (Coco), Happy (from Despicable Me 2), a Star is Born (Shallow), are all arguably wayyyyy more memorable and these are just off the top of my head. We've had quite a few memorable songs in soundtracks this millennium. edit: I had forgotten about Jai Ho (Slumdog Millionaire) and Lose Yourself (8 mile)

It's obviously subjective, but in terms of numbers the only contender in that list is Let It Go, which had about 1/3rd the reach.

Nothing on that list - movies or songs - had the cultural impact of Furious 7 or See You Again.

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#85

I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…

I agree. For some reason the female voices are waaay more convincing than the male ones too, which sound barely better than speech synthesis from a decade ago.

Results correlate to investment, and there’s more in synthesizing female coded voices. As for the why female coded voices gets more investments, we all know, only difference is in attitude towards that (the correct answer, of course, is “it sucks”)

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#87
post #28

I feel like this is a step in the right direction, but a lot of emotive text-to-speech models are only changing the duration and loudness of each word, the timing/pauses are better too. I would love to have a model that can make sense of things like stressing particular syllables or phonemes to make a point.

this model is superb

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#88

I really hope someone within Microsoft is naming their open source coding agent Microsoft VibeCode. Let this be a thing. Its either that or "Lo" then you can have Lo work with Phi, so you can Vibe code with Lo Phi. https://techcommunity.microsoft.com/blog/azure-ai-foundry-bl...

genius

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#89
post #54

I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…

This is close to SOTA emotional performance, at least the female voices. I trust the human scores in the paper. At least my ear aligns with that figure. With stuff like this coming out in the open, I wonder if ElevenLabs will maintain its huge ARR lead in the field. I really don't see how they can continue to maintain a lead when their offering is getting trounced by open models.

11labs is facing a real competitor

Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model

#90

Earlier quoted context omitted.

This is ludicrous. macOS has had text-to-speech for ages with acceptable quality, and they never needed energy- and compute-expensive models for it. And it reacts instantly, not after ridiculous delays. I cannot believe this hype about “AI”, it’s just too absurd.

> with acceptable quality Compared to IBMs Steven Hawking's chair, maybe. But apple tts is not acceptable quality in any modern understanding of SotA, IMO.

Different use cases:

If you need a not-visual output of text, SoyA is a waste of electrons.

If you want to try and mimic a human speaker, then it ain’t.

Question is why would you need to have the computer sound more human, except for “because I can”.

Post reply on HN