VibeVoice: A Frontier Open-Source Text-to-Speech Model
131–140 of 177 posts
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#132I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…
It's good but not the best free model. I find Chatterbox to be more realistic with no robot-sounding and better (though not perfect) intonation.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#133A 100M podcast model
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#134Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#135Earlier quoted context omitted.
> with acceptable quality Compared to IBMs Steven Hawking's chair, maybe. But apple tts is not acceptable quality in any modern understanding of SotA, IMO.
Different use cases: If you need a not-visual output of text, SoyA is a waste of electrons. If you want to try and mimic a human speaker, then it ain’t. Question is why would you need to have the computer sound more human, except for “because I can”.
There's a lot of stuff I don't have time to sit down and read, but want to listen to while I cook/laundry/shower/drive/etc.
Often recordings don't exist. Or when they do, an audiobook just has a bad voiceover artist, or one that just rubs you the wrong way.
The more human text-to-speech sounds, the easier and less distracting it is to listen to. There's real value in it, it's not "because I can".
You know how it's nicer to read in 300 dpi instead of 72 dpi? Or in Garamond rather than Courier? Or in Helvetica rather than Comic Sans? It's like that, only for speech.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#136Earlier quoted context omitted.
We all know? Female voices have better intelligibility? That's my guess anyway.
There's a lot of money and effort spent in satisfying the sexual desires of (predominantly straight) men. There's not typically quite as much interest in doing the same for women. For example I've been looking at models and loras for generating images, and the boards are _full_ of ones that will generate women well or in some particular style. Quite often at least a couple of the preview images for each are hidden be…
Female voices are often rated as being clearer, easier to understand, "warmer", etc.
Why this is the case is still an open question, but it's definitely more complex than just SEX.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#137I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#138MIT license - very nice!
The application of known FOSS licenses to what is effectively a binary-only release is misleading and borderline meaningless.
If you're in a company and need a model which one do you think you're getting past compliance & legal - the one that says MIT or the one that says "non-commercial use only"?
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#139I read the comments praising these voices as very life like, and went to the page primed to hear very convincing voices. That is not at all what I heard though. The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninsp…
Their comments about the singing and background music are odd. It’s been a while since I’ve done academic research, but something about those comments gave me a strong “we couldn’t figure out how to make background music go away in time for our paper submission, so we’re calling it a feature” vibe as opposed to a “we genuinely like this and think its a differentiator” vibe.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#140Earlier quoted context omitted.
Knowing the history of Microsoft marketing, it will either be called something like "Microsoft Copilot Code Generator for VSCode" or something like "Zunega"...
Well don't forget "Microsoft SQL" ;) They'll name something as though they invented it and now have the worse possible way to google it.