Live data from Hacker News

This Voice Doesn't Exist – Generative Voice AI

blog.elevenlabs.io

81–90 of 279 posts

Re: This Voice Doesn't Exist – Generative Voice AI

#81

The “narrative” example is pretty good, but the “conversational” example is rather unpleasant to listen to. (Especially if you know how well Meryl Streep delivers that monologue in the original: https://youtu.be/Ja2fgquYTCg )

Let's talk about this "Narrative" example.

When I listened to it, my first impression was that it must be the real actor they included for comparison purposes but that they failed to label it correctly. I thought it is not machine-generated. I couldn't tell the slightest artifact except what sounded like a low-bitrate sound encoding (maybe using a codec geared toward speech). Can you tell anything "off" about it?

As for the encoding artifact such as a tinny sound or low-bitrate sound, that is the type you hear on an MP3 or low bitrate codec for speech. For example, when I record a message on https://vocaroo.com/ the "premier" voice recording service it sounds 10x worse. Here is a sample I just recorded of my own speech: https://voca.ro/18oSJ1sHU5w5

After my first impression that the narrative example might be a real human mislabelled for comparison purposes, I listened to the next two, labelled News and Conversational. I found these very easy to tell as AI-generated.

Thinking back to why I found the narrative example so compelling, I thought perhaps the issue is that the first example is in British English which I'm less used to than American English. I grew up in the United States. Perhaps since the accent doesn't match my own, it is harder for me to perceive it as generated.

-> Can a native speaker of British English tell us whether listening to the first example you can tell in any way that it is a robot? Maybe it is as obvious to you as the next two are to me.

Still, I've listened to a fair amount of British English in my life so perhaps there is an alternative explanation for why the first one was better. For example, it could have been trained on a reader's voice who has narrated thousands of hours in very high studio quality in a fairly consistent way, leaving this type of text much easier to synthesize than the other two examples due to more training data or higher-quality audio.

For me, the first one is really indistinguishable from a narrator's true voice, though it does sound a bit tinny which could also happen as an artifact of the recording process.

In terms of "how confident are you that this is a real person" the second two examples I would put at 0 - it's totally obvious that it is not a real person, whereas the first one sounds like a 10 to me: obviously a real narrator. (With a bit of artifacting that sounds like an mp3.)

[1] The text is here https://www.nytimes.com/2001/11/19/books/chapters/the-lord-o...

Re: This Voice Doesn't Exist – Generative Voice AI

#82
Interestingly, some of the robot styles take a very obvious and dramatic fake breath. I say "fake" since a robot doesn't need to breathe and it's not exactly considered a phoneme. The fake breaths don't really make the robot sound more convincing.

When you listen to the first example labelled "Narrative" you can tell where a human speaker would have inhaled (which is something the AI could have picked up on from copious training data) though the inhale itself could be muted in post-editing, e.g. after the long 24-word first phrase[1] ending in "special magnificence", and then again at the end of the sentence. It could just be the way the AI reads the comma but it is very convincing.

The "News" and "Conversational" examples don't include that pause effect. In the cerulean monologue, there is no pause after "for instance" despite it being in the monologue.

However, the robot takes a deep dramatic breath after the word "I see"[2]. " Oh, okay. I see, [DEEP LOUD DRAMATIC BREATH BY ROBOT], you think this has nothing to do with you. [LOUD DRAMATIC HALF BREATH BY ROBOT] You go to your closet and you select I don't know that lumpy blue sweater for instance because you're trying to tell the world that you take yourself". There is no pause on the comma around "for instance" though the script has one. I decided to check whether the robot is just copying the original film exactly and that's not it either.[3]

Comparison:

    Robot: "Oh, okay. I see, [DEEP LOUD DRAMATIC BREATH BY ROBOT], you think this has nothing to do with you. [LOUD DRAMATIC HALF BREATH BY ROBOT] You go to your closet [no breath] and you select I don't know that lumpy blue sweater for instance [QUICK HALF BREATH BY ROBOT] because you're trying to tell the world [no breath] that you take yourself too seriously to care about what you put on your back but [no breath] what you don't know is that sweater is not just blue it's not turquoise it's not lapis it's actually cerulean."

    Original: "Oh, okay. I see [no breath] you think this has nothing to do with you. [loud long breath] You go to your closet [breath] and you select I don't know that lumpy blue sweater for instance [no breath] because you're trying to tell the world that you [breath] take yourself too seriously to care about what you put on your back but [breath] what you don't know is that sweater is not just blue it's not turquoise it's not lapis it's actually cerulean."
Text: "Oh, okay. I see, you think this has nothing to do with you.

You… go to your closet, and you select… I don’t know, that lumpy blue sweater for instance, because you’re trying to tell the world that you take yourself too seriously to care about what you put on your back, but what you don’t know is that that sweater is not just blue, it’s not turquoise, it’s not lapis, it’s actually cerulean. "

I've annotated the breaths in the "conversational" robot sample vs the original film:

                     Robot                  Original                Same/different?
     I see...        [Loud breath]          [no breath]             Different
     with you...     [Loud quick breath]    [loud long breath]      Similar
     your closet...  [no breath]            [breath]                Different
     for instance... [QUICK half breath]    [no breath]             Different
     that you...     [no breath]            [breath]                Different
     back but...     [no breath]            [breath]                Different
The robot's loud dramatic breath is unmistakable, but it's clear it's not copying the source exactly, since it occurs at different places.

[1] The text is here: https://www.nytimes.com/2001/11/19/books/chapters/the-lord-o...

[1] The text is here: https://artdepartmental.com/blog/devil-wears-prada-cerulean-...

[2] https://www.youtube.com/watch?v=us52N76XA28&t=1m24s

Re: This Voice Doesn't Exist – Generative Voice AI

#83
By the way if anyone is in this thread due to working on AI speech synthesis for any company, I am interested in AI as well as audio production and I would love to talk about joining the team as an AI researcher. Just send me some mail, my email is in my profile.

Re: This Voice Doesn't Exist – Generative Voice AI

#84
What are the odds of this kind of thing being open source so I can use it at home. So far, most of the "good" text-to-speech systems are all commercial services

https://aws.amazon.com/polly/

https://cloud.google.com/text-to-speech

https://azure.microsoft.com/en-us/products/cognitive-service...

And now one is also a service.

I tried using tortoise-tts on my M1. Generating a 7 minute speech took 3 days and, while better than the 15 yr old text-to-speech built into the OS it wasn't close to the quality of the services above. Maybe I don't know who to use it but of course it's not as simple as text-to-speech. You need the system to ideally understand the text it can act out parts

Of course see my username. I want to generate personal adult content so I'd prefer not to upload it to a service.

Re: This Voice Doesn't Exist – Generative Voice AI

#86
I'd like to see this technology become cheap and ubiquitous enough that everyone can choose for themselves what voice they would like to hear right at the moment of consumption. It's always a huge bummer when there's a book I want to listen to on audible with terrible narration. Somebody must have liked that voice for the person to be hired, but people's tastes differ and sometimes the people they've selected just really grate on my ears.

It would also be cool if celebrities / existing voice talent could somehow license the synthesis of their voice. I read something about James Earl Jones doing this with Disney for future Star Wars projects. I'm sure there are people out there who would love to have every work they listen to be in the voice of their favorite narrator/celebrity.

Re: This Voice Doesn't Exist – Generative Voice AI

#87

I can't tell if I'm starting to get that old person "new things are scary" instinct or if my gut level of fear about the implications of these things is warranted. As impressive as a lot of these models are, I can't help but feel like they're going to end up making an incredible amount of sterile soulless content that makes everyone's lives worse . We're already drowning in ad dominated cynical soulless computer gene…

[deleted]

Re: This Voice Doesn't Exist – Generative Voice AI

#88

I can't tell if I'm starting to get that old person "new things are scary" instinct or if my gut level of fear about the implications of these things is warranted. As impressive as a lot of these models are, I can't help but feel like they're going to end up making an incredible amount of sterile soulless content that makes everyone's lives worse . We're already drowning in ad dominated cynical soulless computer gene…

I think we already have a lot of soulless human generated search results. I think there will be need for a greater level of filtering and curation yes, but I see it as an opportunity both for creators and curators. The barriers to entry for media creation will go down, but with saturation also the already low margins of profit will get worse.

Also, AI will do the filtering, not just blocking uninteresting content, but actually removing known and uninteresting information from content.

Re: This Voice Doesn't Exist – Generative Voice AI

#89

Still sounds pretty fake to me. There’s a hurriedness to the speech and a monotonic uniformity in enunciation that is uncannily machine. Good to know that voice actors will have jobs for a while longer…

I thought the Narrative one was 100% there. I'd still give the News one 99% and Conversational 98%.

Yes, for the sake of humanity, I hope the examples are cherry picked and The lord of the rings audiobook is in the train set...

Re: This Voice Doesn't Exist – Generative Voice AI

#90

I can't tell if I'm starting to get that old person "new things are scary" instinct or if my gut level of fear about the implications of these things is warranted. As impressive as a lot of these models are, I can't help but feel like they're going to end up making an incredible amount of sterile soulless content that makes everyone's lives worse . We're already drowning in ad dominated cynical soulless computer gene…

I agree with this completely. Technology has always made us trade quality for low-quality quantity in exchange for convenience. People now interact more through technology which removes a lot of body language and other enriching experiences. The most dangerous aspect of this is that each step seems relatively harmless: right now, ChatGPT and DALL-E are amusements, but each small step is building a monstrous and as yo…

Why are you on a forum called Hacker News if you hate technology so much?
Post reply on HN