Why do seemingly all these text-to-speech programs attempt to produce spoken voice based solely on raw text? Why don't they consume a MIDI-like text-markup language where you can write phonetic pronunciations along with markup about the emotion, volume, speed, etc.? I feel like this is a huge unnecessary roadblock holding back this kind of technology. It'd be like if every music composition program rendered a wave file not by MIDI or VST, but by trying to visually read sheet music. I totally understand why TTS solutions that have to consume arbitrary content, like screen-readers, need to read purely raw text. But content creators don't need to be limited to raw text! Why is everyone doing it that way? Where is the TTS markup language for content creators?
This Voice Doesn't Exist – Generative Voice AI
111–120 of 279 posts
Re: This Voice Doesn't Exist – Generative Voice AI
#112I can't tell if I'm starting to get that old person "new things are scary" instinct or if my gut level of fear about the implications of these things is warranted. As impressive as a lot of these models are, I can't help but feel like they're going to end up making an incredible amount of sterile soulless content that makes everyone's lives worse . We're already drowning in ad dominated cynical soulless computer gene…
I agree with this completely. Technology has always made us trade quality for low-quality quantity in exchange for convenience. People now interact more through technology which removes a lot of body language and other enriching experiences. The most dangerous aspect of this is that each step seems relatively harmless: right now, ChatGPT and DALL-E are amusements, but each small step is building a monstrous and as yo…
It is perfectly viable in the modern day, to work a job, have passionate hobbies, regularly meet for social events, volunteer, etc and spend minimal to zero time engaging on the internet, besides pragmatic things like map directions
Re: This Voice Doesn't Exist – Generative Voice AI
#113Earlier quoted context omitted.
> Technology has always made us trade quality for low-quality quantity in exchange for convenience. People now interact more through technology which removes a lot of body language and other enriching experiences. I went to the mall today and you can tell malls are dying. I lived in a small town where the mall died and it had a zombie like existence a long time before it finally cratered. The mall here in this larger…
I never went to malls to socialize.
Re: This Voice Doesn't Exist – Generative Voice AI
#114This is cooler than ChatGPT and image generation as far as I'm concerned. If they're able to bring out the emotional connectivity and purposefulness of the human voice, it will be revolutionary...
There are so many uses cases for this, even with the current quality. Many game developers dream of having something like this.
Re: This Voice Doesn't Exist – Generative Voice AI
#115Earlier quoted context omitted.
Not only voice actors, include radio hosts, documentary/news content, any voice over for anything, as well as imitation of familiar voices.
This will really open pandoras box for scammers and other bad actors. Grandma won't know she's speaking with an AI.
spotting open (closed) AI models by doing the Voight-KAPTCHA test
Re: This Voice Doesn't Exist – Generative Voice AI
#116Okay can I ask a question that has been bothering me for a long time? Why do seemingly all these text-to-speech programs attempt to produce spoken voice based solely on raw text ? Why don't they consume a MIDI-like text-markup language where you can write phonetic pronunciations along with markup about the emotion, volume, speed, etc.? I feel like this is a huge unnecessary roadblock holding back this kind of technol…
https://en.wikipedia.org/wiki/Speech_Synthesis_Markup_Langua...
Amazon Polly, (which seems kind of ancient with all these new solutions showing up) has supported SSML for some time.
AWS Polly SSML docs: https://docs.aws.amazon.com/polly/latest/dg/ssml.html
Re: This Voice Doesn't Exist – Generative Voice AI
#117Earlier quoted context omitted.
I work in TTS and i just dont believe this. If these really are random text and not trained on literally the copy they are reading, with no correction I would be surprised. Also, our competitors have good voices but they also take ages to produce. Maybe these really are legit but take like 1 minute to produce or something. So while this is impressive, i doubt that in practice this would be this high quality and could…
I want to agree, but I searched on their website and found their narration service with 2 full book examples. I listened to the first one for a while and it's the first time an Ai narrator was good enough to keep me listening: https://www.audiostory.ai/2065785/11707800-alice-s-adventure...
Re: This Voice Doesn't Exist – Generative Voice AI
#118I can't tell if I'm starting to get that old person "new things are scary" instinct or if my gut level of fear about the implications of these things is warranted. As impressive as a lot of these models are, I can't help but feel like they're going to end up making an incredible amount of sterile soulless content that makes everyone's lives worse . We're already drowning in ad dominated cynical soulless computer gene…
Re: This Voice Doesn't Exist – Generative Voice AI
#119Is that even a thing? You can't copyright a voice. There can be a personality right under state law, but the main case on that was someone hired to sound like Bette Midler for a commercial.
Re: This Voice Doesn't Exist – Generative Voice AI
#120Okay can I ask a question that has been bothering me for a long time? Why do seemingly all these text-to-speech programs attempt to produce spoken voice based solely on raw text ? Why don't they consume a MIDI-like text-markup language where you can write phonetic pronunciations along with markup about the emotion, volume, speed, etc.? I feel like this is a huge unnecessary roadblock holding back this kind of technol…