Live data from Hacker News

This Voice Doesn't Exist – Generative Voice AI

blog.elevenlabs.io

131–140 of 279 posts

Re: This Voice Doesn't Exist – Generative Voice AI

#132
post #103

Interestingly, some of the robot styles take a very obvious and dramatic fake breath. I say "fake" since a robot doesn't need to breathe and it's not exactly considered a phoneme. The fake breaths don't really make the robot sound more convincing. When you listen to the first example labelled "Narrative" you can tell where a human speaker would have inhaled (which is something the AI could have picked up on from copi…

I noticed a breath in the demo audio in the linked article and while it stood out, I was impressed by it rather than thinking it felt forced. I'm sure if I listened to enough AI voice it would stand out more and feel forced.

Did you find the whole clip it was in convincing? For me, I didn't even notice the breath but the entire second and third clip felt obviously AI-generated. But the first clip sounded absolutely real (maybe with some compression artifacts - see my other comment.)

Later when I went back and listened carefully for why the first clip felt so "real" I noticed it had pauses. (No breaths per se but they are sometimes removed from edited audio.) However, I then noticed that the conversational clip, which felt unnatural to me, had very obvious breaths. The entire effect of the conversational clip didn't sound like a human at all. It sounded like an AI.

Did you find the whole conversational clip "convincing"? (Did it sound like a human to you?) How about the narration clip?

Re: This Voice Doesn't Exist – Generative Voice AI

#133
post #111

Okay can I ask a question that has been bothering me for a long time? Why do seemingly all these text-to-speech programs attempt to produce spoken voice based solely on raw text ? Why don't they consume a MIDI-like text-markup language where you can write phonetic pronunciations along with markup about the emotion, volume, speed, etc.? I feel like this is a huge unnecessary roadblock holding back this kind of technol…

There's a pretty advanced Mac OS speech markup language, I wrote about it here: https://www.mattmontag.com/personal/mac-os-x-speech-synthesi...

Going back further, there was also a prosody markup for Sound Blaster speech synthesis (Dr. Sbaitso, anyone?).

Re: This Voice Doesn't Exist – Generative Voice AI

#134
post #99

Earlier quoted context omitted.

I agree with this completely. Technology has always made us trade quality for low-quality quantity in exchange for convenience. People now interact more through technology which removes a lot of body language and other enriching experiences. The most dangerous aspect of this is that each step seems relatively harmless: right now, ChatGPT and DALL-E are amusements, but each small step is building a monstrous and as yo…

> Technology has always made us trade quality for low-quality quantity in exchange for convenience. People now interact more through technology which removes a lot of body language and other enriching experiences. I went to the mall today and you can tell malls are dying. I lived in a small town where the mall died and it had a zombie like existence a long time before it finally cratered. The mall here in this larger…

Or don't shop at all and use that extra time to walk with friends in nature. Or when you really do need to shop, avoid the commute and use that extra time to spend with friends in nature. Being forced to be around strangers to get chores done doesn't put soul into my life.

I also avoid laundromats and do laundry at home and it doesn't feel soulless.

Re: This Voice Doesn't Exist – Generative Voice AI

#135
I always wondered why those generative voices dont capture the feeling of the text per segment and incorporate it to the output e.g. news, narration, first person hunted by vampires, whatever. Seems like a kind of low hanging fruit.

Disclaimer: I use tons of audiobooks so that might not be what people need in general

Re: This Voice Doesn't Exist – Generative Voice AI

#137

I can't tell if I'm starting to get that old person "new things are scary" instinct or if my gut level of fear about the implications of these things is warranted. As impressive as a lot of these models are, I can't help but feel like they're going to end up making an incredible amount of sterile soulless content that makes everyone's lives worse . We're already drowning in ad dominated cynical soulless computer gene…

I agree with this completely. Technology has always made us trade quality for low-quality quantity in exchange for convenience. People now interact more through technology which removes a lot of body language and other enriching experiences. The most dangerous aspect of this is that each step seems relatively harmless: right now, ChatGPT and DALL-E are amusements, but each small step is building a monstrous and as yo…

And mind-boggling you criticize and demonize technology on a forum about technology all while not only being on the internet but also using electricity, a computer, and certainty surrounded by gadgets and other amenities of modern life.

Nature is SHIT, that is why people created technology. There is nothing preventing you from going to the middle of nowhere and reject modernity, no one is forcing you, but you are that because you wanted and liked it. You talk people should have an "instinctual revulsion" towards technology, but not even you yourself has this reaction towards technology because it is a stupid idea that not even luddites like you commit to it.

If anything the technology we have nowadays is not even 0,01% of what we should have. We should have the technology to make any movie anyone ever wanted to see in a blink of an eye, all done in the best quality ever imagined. We should have the power to build a Dyson Sphere around the sun to harness its energy. We should be able to construct fully immersive virtual reality, like San Junipero from the Black Mirror's episode, we should have the power extend human life indefinitely.

Re: This Voice Doesn't Exist – Generative Voice AI

#138
post #111

Okay can I ask a question that has been bothering me for a long time? Why do seemingly all these text-to-speech programs attempt to produce spoken voice based solely on raw text ? Why don't they consume a MIDI-like text-markup language where you can write phonetic pronunciations along with markup about the emotion, volume, speed, etc.? I feel like this is a huge unnecessary roadblock holding back this kind of technol…

As a note, there are indeed markup formats to write the phonetic pronunciations, and also allowing everything you mentioned.

It's called SSML.

Re: This Voice Doesn't Exist – Generative Voice AI

#139
post #111

Okay can I ask a question that has been bothering me for a long time? Why do seemingly all these text-to-speech programs attempt to produce spoken voice based solely on raw text ? Why don't they consume a MIDI-like text-markup language where you can write phonetic pronunciations along with markup about the emotion, volume, speed, etc.? I feel like this is a huge unnecessary roadblock holding back this kind of technol…

There's lots of text and audio already without this, that's probably the key factor practically. Similarly then for use cases, converting text that already exists is much more approachable than creating new marked up text.

Tortoise lets you add prompts into the text like [I am angry] which modifies the voice interestingly.

Re: This Voice Doesn't Exist – Generative Voice AI

#140
About a month ago, I made a toy bot that listens to your voice with OpenAI Whisper, generates a response with GPT-2 and vocalizes the response using the Eleven Labs. The TTS quality produced by the Eleven Labs algorithm was mind-blowing to me. The API that they provided was super easy to use. Good product!
Post reply on HN