VibeVoice: A Frontier Open-Source Text-to-Speech Model
81–90 of 177 posts
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#82I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#83Earlier quoted context omitted.
The most popular song of the year from one of the most popular movie franchises that had been in the global news due to the death of its star. Probably the most memorable song from a soundtrack of the century so far.
I'm Just Ken (Barbie), Skyfall, Let it Go (Frozen), Remember Me (Coco), Happy (from Despicable Me 2), a Star is Born (Shallow), are all arguably wayyyyy more memorable and these are just off the top of my head. We've had quite a few memorable songs in soundtracks this millennium. edit: I had forgotten about Jai Ho (Slumdog Millionaire) and Lose Yourself (8 mile)
Nothing on that list - movies or songs - had the cultural impact of Furious 7 or See You Again.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#84Is there a current, updated list (ideally, a ranking) of the best open weights TTS models? I'm actually more interested in STT (ASR) but the choices there are rather limited.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#85I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…
I agree. For some reason the female voices are waaay more convincing than the male ones too, which sound barely better than speech synthesis from a decade ago.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#86MIT license - very nice!
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#87I feel like this is a step in the right direction, but a lot of emotive text-to-speech models are only changing the duration and loudness of each word, the timing/pauses are better too. I would love to have a model that can make sense of things like stressing particular syllables or phonemes to make a point.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#88I really hope someone within Microsoft is naming their open source coding agent Microsoft VibeCode. Let this be a thing. Its either that or "Lo" then you can have Lo work with Phi, so you can Vibe code with Lo Phi. https://techcommunity.microsoft.com/blog/azure-ai-foundry-bl...
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#89I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…
This is close to SOTA emotional performance, at least the female voices. I trust the human scores in the paper. At least my ear aligns with that figure. With stuff like this coming out in the open, I wonder if ElevenLabs will maintain its huge ARR lead in the field. I really don't see how they can continue to maintain a lead when their offering is getting trounced by open models.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#90Earlier quoted context omitted.
This is ludicrous. macOS has had text-to-speech for ages with acceptable quality, and they never needed energy- and compute-expensive models for it. And it reacts instantly, not after ridiculous delays. I cannot believe this hype about “AI”, it’s just too absurd.
> with acceptable quality Compared to IBMs Steven Hawking's chair, maybe. But apple tts is not acceptable quality in any modern understanding of SotA, IMO.
If you need a not-visual output of text, SoyA is a waste of electrons.
If you want to try and mimic a human speaker, then it ain’t.
Question is why would you need to have the computer sound more human, except for “because I can”.