[deleted - I'm an idiot]
Whisper is speech-to-text. VibeVoice is text-to-speech.
VibeVoice: A Frontier Open-Source Text-to-Speech Model
21–30 of 177 posts
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#22So that's a useful next step: for multi-voice TTS models, make them sound like they're in the same room.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#23The voices are decent, but the intonation is off on almost every phrase, and there is a very clear robotic-sounding modulation. It's generally very impressive compared to many text-to-speech solutions from a few years ago, but for today, I find it very uninspiring. The AI generated voice you hear all over YouTube shorts is at least as good as most of the samples on this page.
The only part that seemed impressive to me was the English + (Mandarin?) Chinese sample, that one seemed to switch very seamlessly between the two. But this may well be simply because (1) I'm not familiar with any Chinese language, so I couldn't really judge the pronunciation of that, and (2) the different character systems make it extremely clear that the model needs to switch between different languages. Peut-être que cela n'aurait pas été si simple if it had been switching between two languages using the same writing system - I'm particularly curious how it would have read "simple" in the phrase above (I think it should be read with the French pronunication, for example).
And, of course, the singing part is painfully bad, I am very curious why they even included it.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#24Wow. I admit that I am not a native speaker, but this looks (or rather, sounds) VERY impressive and I could mistake it for hearing two people talking.
I generally don't like a lot of the AI generated slop that's starting to pop up on YouTube these days... I do enjoy some of the reddit story channels, but have completely stopped with it all now. With the AI stuff, it really becomes apparent with dates/ages and when numbers are spoken. Dates/ages/timelines are just off as far as story generation, and really should be human tweaked. As to the voice gen, saying a year or measurement is just not how English speakers (US or otherwise) speak.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#25I really hope someone within Microsoft is naming their open source coding agent Microsoft VibeCode. Let this be a thing. Its either that or "Lo" then you can have Lo work with Phi, so you can Vibe code with Lo Phi. https://techcommunity.microsoft.com/blog/azure-ai-foundry-bl...
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#26Wow. I admit that I am not a native speaker, but this looks (or rather, sounds) VERY impressive and I could mistake it for hearing two people talking.
The giveaway is they will never talk over each other. Only one speaker at a time, consistently.
Would probably want to do similar to balance crossfade anyway... having each speaker's input offset from center instead of straight mono.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#27I really hope someone within Microsoft is naming their open source coding agent Microsoft VibeCode. Let this be a thing. Its either that or "Lo" then you can have Lo work with Phi, so you can Vibe code with Lo Phi. https://techcommunity.microsoft.com/blog/azure-ai-foundry-bl...
Knowing the history of Microsoft marketing, it will either be called something like "Microsoft Copilot Code Generator for VSCode" or something like "Zunega"...
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#28I would love to have a model that can make sense of things like stressing particular syllables or phonemes to make a point.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#29Earlier quoted context omitted.
Knowing the history of Microsoft marketing, it will either be called something like "Microsoft Copilot Code Generator for VSCode" or something like "Zunega"...
Well don't forget "Microsoft SQL" ;) They'll name something as though they invented it and now have the worse possible way to google it.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#30I'm actually more interested in STT (ASR) but the choices there are rather limited.