Earlier quoted context omitted.
Results correlate to investment, and there’s more in synthesizing female coded voices. As for the why female coded voices gets more investments, we all know, only difference is in attitude towards that (the correct answer, of course, is “it sucks”)
We all know? Female voices have better intelligibility? That's my guess anyway.
VibeVoice: A Frontier Open-Source Text-to-Speech Model
111–120 of 177 posts
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#112Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#113Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#114I really hope someone within Microsoft is naming their open source coding agent Microsoft VibeCode. Let this be a thing. Its either that or "Lo" then you can have Lo work with Phi, so you can Vibe code with Lo Phi. https://techcommunity.microsoft.com/blog/azure-ai-foundry-bl...
Knowing the history of Microsoft marketing, it will either be called something like "Microsoft Copilot Code Generator for VSCode" or something like "Zunega"...
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#115Looking forward to the day when tts and speech recognition will work on Croatian, or other less prevalent languages. It seems that it's only variants of English, Spanish and Chinese which are somewhat working.
Have you tried Soniox for speech recognition? It supports Croatian. Or are you just looking for self-hosted open-source models? Soniox is very cheap ($0.1/h for async, $0.12/h for real-time) and you get $200 free credits on signup. https://soniox.com/ Disclaimer: I used to work for Soniox
In Android Auto / CarPlay I can't even get voice guidance that works properly, much less reading notifications, or composing a reply using STT
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#116Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#117What an odd name to me, becaus "Vibe" is, in my mind, equal to somewhat poor quality. Like "Vibe Coding". But that's probably just some bias from my side.
Vibe coding just became a term this spring. I doubt that that the substantial part, like giving it a project code name and getting company approval of this research project started after that. It's not libe vibe has a negative connotation in general yet.
But I do agree with you in that generally there's probably no negative connotation (yet).
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#118Earlier quoted context omitted.
Results correlate to investment, and there’s more in synthesizing female coded voices. As for the why female coded voices gets more investments, we all know, only difference is in attitude towards that (the correct answer, of course, is “it sucks”)
We all know? Female voices have better intelligibility? That's my guess anyway.
For example I've been looking at models and loras for generating images, and the boards are _full_ of ones that will generate women well or in some particular style. Quite often at least a couple of the preview images for each are hidden behind a button because they contain nudity. Clearly the intent is that they are at least able to generate porn containing women. There's a small handful that are focused on men and they're very aware of it, they all have notes lampshading how oddball they are to even exist.
I would expect that this is not as pronounced an effect in the world generating speech, but it must still exist.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#119Earlier quoted context omitted.
> with acceptable quality Compared to IBMs Steven Hawking's chair, maybe. But apple tts is not acceptable quality in any modern understanding of SotA, IMO.
Different use cases: If you need a not-visual output of text, SoyA is a waste of electrons. If you want to try and mimic a human speaker, then it ain’t. Question is why would you need to have the computer sound more human, except for “because I can”.
I think translation would be a big use - maybe translating your voice to another language while maintaining emotion and intonation, or dubbing content (videos, movies, podcasts, ...) that isn't otherwise available in your native language.
Traditional non-ML TTS for longer content like podcasts or audiobooks seems like it'd become grating to the point of being unlistenable, or at least a significantly worse experience. Stands to benefit from more natural sounding voices that can place emphasis in the right places.
Since Stephen Hawking was brought up, there are likely also people with voice-impairing illnesses who would like to speak in their own voice again (in addition to those who are fine with a robotic voice). Or alternatively, people who are uncomfortable with their natural voice and want to communicate closer to how they wish to be perceived.
Could also potentially be used for new forms of interactive media that aren't currently feasible - customised movies, audio dramas where the listener plays a role, videogame NPCs that react with more than just prerecorded lines, etc.
Re: VibeVoice: A Frontier Open-Source Text-to-Speech Model
#120seemingly supports only English, Indian and Chinese