Live data from Hacker News

AI Clones Your Voice After Listening for 5 Seconds (2018)

google.github.io

51–60 of 338 posts

Re: AI Clones Your Voice After Listening for 5 Seconds (2018)

#51
So....I'm going to paste the abstract here because the headline is incredibly misleading and should be changed.

>Abstract: We describe a neural network-based system for text-to-speech (TTS) synthesis that is able to generate speech audio in the voice of many different speakers, including those unseen during training. Our system consists of three independently trained components: (1) a speaker encoder network, trained on a speaker verification task using an independent dataset of noisy speech from thousands of speakers without transcripts, to generate a fixed-dimensional embedding vector from seconds of reference speech from a target speaker; (2) a sequence-to-sequence synthesis network based on Tacotron 2, which generates a mel spectrogram from text, conditioned on the speaker embedding; (3) an auto-regressive WaveNet-based vocoder that converts the mel spectrogram into a sequence of time domain waveform samples. We demonstrate that the proposed model is able to transfer the knowledge of speaker variability learned by the discriminatively-trained speaker encoder to the new task, and is able to synthesize natural speech from speakers that were not seen during training. We quantify the importance of training the speaker encoder on a large and diverse speaker set in order to obtain the best generalization performance. Finally, we show that randomly sampled speaker embeddings can be used to synthesize speech in the voice of novel speakers dissimilar from those used in training, indicating that the model has learned a high quality speaker representation.

Re: AI Clones Your Voice After Listening for 5 Seconds (2018)

#55

So....I'm going to paste the abstract here because the headline is incredibly misleading and should be changed. >Abstract: We describe a neural network-based system for text-to-speech (TTS) synthesis that is able to generate speech audio in the voice of many different speakers, including those unseen during training. Our system consists of three independently trained components: (1) a speaker encoder network, trained…

Do you want to elaborate on how the title is misleading? From reading the abstract it seems accurate to me.

Re: AI Clones Your Voice After Listening for 5 Seconds (2018)

#56

The malign applications of this technology greatly outweigh the benign. Discuss.

Con/Pro, depending on perspective: people will have to give up the illusion that they ever really could definitively tell truth from fiction.

Re: AI Clones Your Voice After Listening for 5 Seconds (2018)

#58

Wow, impressive results! Already a few examples in the comments of what bad actors could do this tech. I wanted to share an example of something good. I lost my dad about 6 years ago after a Stage 4 cancer diagnosis and a 3 month rapid diagnosis. I have some, but not a lot of video content of him from over the years. My mom still misses him terribly so for her 60th birthday I tried to splice together an audio message…

I'm deeply sorry for your loss. Thanks for sharing your story.

Re: AI Clones Your Voice After Listening for 5 Seconds (2018)

#59

The malign applications of this technology greatly outweigh the benign. Discuss.

On a benign level, many VO artists will find themselves out of work now that we can have Don LaFontaine back. On more positive outlook, perhaps this, along with deepfakes, propels us faster towards an evidence-based society.

On a related note, I can definitely foresee a lot of voice actors having their voices cloned for uses they wouldn't really intend. Seems like a big legal grey area as many countries have personality rights.

Re: AI Clones Your Voice After Listening for 5 Seconds (2018)

#60
Saw this on Minute papers last night and had a discussion with my partner about if we needed a secret password or not to tell if it were really one or the other on the phone. I figured that we had enough shared history that that wouldn't be a problem. Then we realized that there's no such thing as a simulated sense of humor yet and that that would be the best natural encryption to any communication.
Post reply on HN