VibeVoice: A Frontier Open-Source Text-to-Speech Model
51–60 of 177 posts
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#52Someone else mentioned in this thread that you cannot add annotations to the text to control the output. I think for these models to really level up there will have to be an intermediate step that takes your regular text as input and it generates an annotated output, which can be passed to the TTS model. That would give users way more control over the final output, since they would be able to inspect and tweak any details instead of expecting the model to get everything correctly in a single pass.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#53Ah, yes, the Furious 7 soundtrack. Definitely something everyone recalls
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#54I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…
I trust the human scores in the paper. At least my ear aligns with that figure.
With stuff like this coming out in the open, I wonder if ElevenLabs will maintain its huge ARR lead in the field. I really don't see how they can continue to maintain a lead when their offering is getting trounced by open models.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#55Some of them have tone wobbles which iirc was more common in early TTS models. Looks like the huge context window is really helping out here.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#56Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#57I tried some TTS models a while ago, but I noticed that none of them allowed to put markup statements in the text. For example, it would be nice to do something like: Hey look! [enthusiastic] Should we tell the others? Maybe not ... [giggles] etc. In fact, I think this kind of thing is absolutely necessary if you want to use this to replace a voice actor.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#58The Spontaneous Emotion dailog sounds like a team member venting through LLMs. They could have skipped the singing part, it would be better if the model did not try to do that :)
1. https://music.youtube.com/watch?v=xl8thVrlvjI&si=dU6aIJIPWSs...
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#59Wow. I admit that I am not a native speaker, but this looks (or rather, sounds) VERY impressive and I could mistake it for hearing two people talking.
The giveaway is they will never talk over each other. Only one speaker at a time, consistently.