VibeVoice: A Frontier Open-Source Text-to-Speech Model
91–100 of 177 posts
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#92I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…
One of the things this model is actually quite good at is voice cloning. Drop a recorded sample of your voice into the voices folder, and it just works.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#93I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…
ElevenLabs has a much more convincing voice model
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#94I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…
Is there any better model you can point at? I would be interested in having a listen. There are people – and it does not matter what it's about – that will overstate the progress made (and others will understate it, case in point). Neither should put a damper on progress. This is the best I personally have heard so far, but I certainly might have missed something.
https://github.com/mlang/llm-tts
Strictly speaking, even music generation fits the usage pattern: text in, audio out.
llm-tts is far from complete, but it makes it relatively "easy" to try a few models in an uniform way.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#95Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#96Earlier quoted context omitted.
> with acceptable quality Compared to IBMs Steven Hawking's chair, maybe. But apple tts is not acceptable quality in any modern understanding of SotA, IMO.
Different use cases: If you need a not-visual output of text, SoyA is a waste of electrons. If you want to try and mimic a human speaker, then it ain’t. Question is why would you need to have the computer sound more human, except for “because I can”.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#97Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#98Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#99they vibecoded their demo website? the text is invisible on Firefox.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#100Bots should never sing.