Hey - developers behind ElevenLabs here. Thank you so much for the constructive and positive feedback - we’re taking it onboard! We’re currently focused on researching and deploying a different way for speech synthesis that can generate nuanced intonation and emotions by understanding text and taking context into account. Additionally, we provide creators with a way to clone their own voice based on very short sample…
This Voice Doesn't Exist – Generative Voice AI
251–260 of 279 posts
Re: This Voice Doesn't Exist – Generative Voice AI
#252Re: This Voice Doesn't Exist – Generative Voice AI
#253Why someone should listen a voice when is faster to read the blog post ?
Re: This Voice Doesn't Exist – Generative Voice AI
#254Earlier quoted context omitted.
I feel like if Bethesda really wants another industry defining game, this is the path they should be taking. AI generated conversation with AI generated voice acting with voice-to-text recognition. You can literally have microphone-voice conversations with NPCs that have rich, AI generated backgrounds and personalities.
Even bigger than that (I think at least) is the potential for fully voiced mods. There’s nothing stopping modders at that point from adding content indistinguishable from the base game.
Re: This Voice Doesn't Exist – Generative Voice AI
#255Earlier quoted context omitted.
I feel like if Bethesda really wants another industry defining game, this is the path they should be taking. AI generated conversation with AI generated voice acting with voice-to-text recognition. You can literally have microphone-voice conversations with NPCs that have rich, AI generated backgrounds and personalities.
How would the union take to that, though? This is not meant as an anti-union comment. I'd just be really surprised if Bethesda ever got to work with union VAs ever again if they went all in on an all-AI voiced game.
You do raise an interesting point, though. Compensation for AI-derived transformations of your work (in this case, your voice) needs to be a thing.
Re: This Voice Doesn't Exist – Generative Voice AI
#256Earlier quoted context omitted.
It's noticable worse than the examples in the blog post. I mean, it's good enough for listening, but no better than the competition.
It's vastly better than any TTS system I have used, but then I've only used a few (mainly phone assistants and the thing built into kindle). What is the competition that you are referring to?
Re: This Voice Doesn't Exist – Generative Voice AI
#257Okay can I ask a question that has been bothering me for a long time? Why do seemingly all these text-to-speech programs attempt to produce spoken voice based solely on raw text ? Why don't they consume a MIDI-like text-markup language where you can write phonetic pronunciations along with markup about the emotion, volume, speed, etc.? I feel like this is a huge unnecessary roadblock holding back this kind of technol…
> I feel like this is a huge unnecessary roadblock holding back this kind of technology. There are speech synthesis markup languages, like SSML. And targeting even lower level has always been possible with commercial speech engines. Think about how tedious and time consuming it is to mark up a large amount of copy? Unless we’re talking about little hints here and there (which is also doable) it rapidly becomes more c…
The first is being able to correct a few things that sound off, as another poster pointed out. "Hey, that's not actually how you pronounce 'synecdoche', it should be 'sɪˈnɛk.doʊ.k'." Or "Less emphasis on the first word, more on the second". Little corrections like that. I imagine a two-stage process where the first generates 'best guess' SSML (or whatever markup) based on the text. Then the content creator can modify it as necessary before it goes into the second step of actual voice synthesis.
The second sweet spot is when your text is dynamically generated. Marking up the entire copy might be a lot of work for pre-written text, but it's a great option for dynamically generated text.
Re: This Voice Doesn't Exist – Generative Voice AI
#258Impressive. Any chance there will be an API version of this product for real-time apps?
Hey - ElevenLabs dev here. The quality above works with <1s latency that for some real-time apps is already sufficient. On smaller chunks of text it can be as quick as ~500ms.
Re: This Voice Doesn't Exist – Generative Voice AI
#259Earlier quoted context omitted.
ElevenLabs dev here - we believe this is a 2 step process and agree it is needed! First, we want to the quality you get out-of-the-box to already by brilliant by taking context into account. Granted, that gets you sometimes 98% there and are working to add manipulation possibility to get you to that 100%; for long-texts though the quality you get is great. For second part, currently TTS providers give complicated tog…
I feel like it would be much harder to create a set of hard controls, like MIDI, to affect the voice acting vs. trying to do a co-embedding space of voices and descriptions of the voices and just saying "Say this quietly and meanly". Thoughts?
Re: This Voice Doesn't Exist – Generative Voice AI
#260Hey - developers behind ElevenLabs here. Thank you so much for the constructive and positive feedback - we’re taking it onboard! We’re currently focused on researching and deploying a different way for speech synthesis that can generate nuanced intonation and emotions by understanding text and taking context into account. Additionally, we provide creators with a way to clone their own voice based on very short sample…
What is your business model? How are you deciding who gets Beta access? What does the voice generation interface look like?
Currently testing Beta with a range of storytelling and publishing use-cases, tackle relevant feedback and make sure the infrastructure supports it. We are planning to open up Beta to everyone by end of this month.
Voice Design interface is currently set of sliders and toggles but currently iterating on what is most accessible.