Live data from Hacker News

This Voice Doesn't Exist – Generative Voice AI

blog.elevenlabs.io

251–260 of 279 posts

Re: This Voice Doesn't Exist – Generative Voice AI

#251

Hey - developers behind ElevenLabs here. Thank you so much for the constructive and positive feedback - we’re taking it onboard! We’re currently focused on researching and deploying a different way for speech synthesis that can generate nuanced intonation and emotions by understanding text and taking context into account. Additionally, we provide creators with a way to clone their own voice based on very short sample…

What is your business model? How are you deciding who gets Beta access? What does the voice generation interface look like?

Re: This Voice Doesn't Exist – Generative Voice AI

#253

Why someone should listen a voice when is faster to read the blog post ?

Absorbing information through your audio input while the visual input is busy with a mindless task is amazing. You can listen to an article or an audiobook while doing laundry.

Re: This Voice Doesn't Exist – Generative Voice AI

#254
post #91

Earlier quoted context omitted.

I feel like if Bethesda really wants another industry defining game, this is the path they should be taking. AI generated conversation with AI generated voice acting with voice-to-text recognition. You can literally have microphone-voice conversations with NPCs that have rich, AI generated backgrounds and personalities.

Even bigger than that (I think at least) is the potential for fully voiced mods. There’s nothing stopping modders at that point from adding content indistinguishable from the base game.

I doubt Bethesda would facilitate this. They'd likely use voice actors to train the voice, and having a famous voice actor saying saucy kink/bdsm/violent things that you tend to see in some mods wouldn't be great PR

Re: This Voice Doesn't Exist – Generative Voice AI

#255
post #91

Earlier quoted context omitted.

I feel like if Bethesda really wants another industry defining game, this is the path they should be taking. AI generated conversation with AI generated voice acting with voice-to-text recognition. You can literally have microphone-voice conversations with NPCs that have rich, AI generated backgrounds and personalities.

How would the union take to that, though? This is not meant as an anti-union comment. I'd just be really surprised if Bethesda ever got to work with union VAs ever again if they went all in on an all-AI voiced game.

I could see Bethesda still wanting to select voice actors to train the model with, especially if they're household names. So they'd get paid.

You do raise an interesting point, though. Compensation for AI-derived transformations of your work (in this case, your voice) needs to be a thing.

Re: This Voice Doesn't Exist – Generative Voice AI

#256

Earlier quoted context omitted.

It's noticable worse than the examples in the blog post. I mean, it's good enough for listening, but no better than the competition.

It's vastly better than any TTS system I have used, but then I've only used a few (mainly phone assistants and the thing built into kindle). What is the competition that you are referring to?

Yeah, as I mentioned I work in TTS and agree with you. If this is legit it is pretty amazing. Certainly would put them as one of the top providers especially given that they could ramp up voice selection. Also, if they truly are training on random stuff they would not have to pay royalties to voice actors since these voices don't exist. This is on par or better then most competitors i am aware of.

Re: This Voice Doesn't Exist – Generative Voice AI

#257
post #111

Okay can I ask a question that has been bothering me for a long time? Why do seemingly all these text-to-speech programs attempt to produce spoken voice based solely on raw text ? Why don't they consume a MIDI-like text-markup language where you can write phonetic pronunciations along with markup about the emotion, volume, speed, etc.? I feel like this is a huge unnecessary roadblock holding back this kind of technol…

> I feel like this is a huge unnecessary roadblock holding back this kind of technology. There are speech synthesis markup languages, like SSML. And targeting even lower level has always been possible with commercial speech engines. Think about how tedious and time consuming it is to mark up a large amount of copy? Unless we’re talking about little hints here and there (which is also doable) it rapidly becomes more c…

I think there are two "sweet spots" here.

The first is being able to correct a few things that sound off, as another poster pointed out. "Hey, that's not actually how you pronounce 'synecdoche', it should be 'sɪˈnɛk.doʊ.k'." Or "Less emphasis on the first word, more on the second". Little corrections like that. I imagine a two-stage process where the first generates 'best guess' SSML (or whatever markup) based on the text. Then the content creator can modify it as necessary before it goes into the second step of actual voice synthesis.

The second sweet spot is when your text is dynamically generated. Marking up the entire copy might be a lot of work for pre-written text, but it's a great option for dynamically generated text.

Re: This Voice Doesn't Exist – Generative Voice AI

#258
post #4

Impressive. Any chance there will be an API version of this product for real-time apps?

Hey - ElevenLabs dev here. The quality above works with <1s latency that for some real-time apps is already sufficient. On smaller chunks of text it can be as quick as ~500ms.

Awesome, thanks for adding that extra color!

Re: This Voice Doesn't Exist – Generative Voice AI

#259

Earlier quoted context omitted.

ElevenLabs dev here - we believe this is a 2 step process and agree it is needed! First, we want to the quality you get out-of-the-box to already by brilliant by taking context into account. Granted, that gets you sometimes 98% there and are working to add manipulation possibility to get you to that 100%; for long-texts though the quality you get is great. For second part, currently TTS providers give complicated tog…

I feel like it would be much harder to create a set of hard controls, like MIDI, to affect the voice acting vs. trying to do a co-embedding space of voices and descriptions of the voices and just saying "Say this quietly and meanly". Thoughts?

Exactly! Only issue is having a well-labelled dataset with those type of cues. We have an idea on how to do it though!

Re: This Voice Doesn't Exist – Generative Voice AI

#260

Hey - developers behind ElevenLabs here. Thank you so much for the constructive and positive feedback - we’re taking it onboard! We’re currently focused on researching and deploying a different way for speech synthesis that can generate nuanced intonation and emotions by understanding text and taking context into account. Additionally, we provide creators with a way to clone their own voice based on very short sample…

What is your business model? How are you deciding who gets Beta access? What does the voice generation interface look like?

We are offering both Speech Synthesis (/TTS) and Voice Lab (Rapid Voice Cloning and Voice Design) as a standard SaaS model (w/ fixed quota of characters you can voice per month). API is directly available on the platform. Outside of standard package that flips to usage-based model and we do tailored deals for custom needs and discounts for high-volume usage.

Currently testing Beta with a range of storytelling and publishing use-cases, tackle relevant feedback and make sure the infrastructure supports it. We are planning to open up Beta to everyone by end of this month.

Voice Design interface is currently set of sliders and toggles but currently iterating on what is most accessible.

Post reply on HN