Live data from Hacker News

Lyrebird – An API to copy the voice of anyone

lyrebird.ai

251–260 of 311 posts

Re: Lyrebird – An API to copy the voice of anyone

#251

Earlier quoted context omitted.

I think Hawking is now so firmly tied to that voice that he would probably never switch for public speaking engagements and the like. I could see him doing such a switch for personal interactions.

I recall his biggest qualm with his current synthesized voice is that it did not come with a British accent. :-) I'm not sure how sentimental he is but he does seem quite tied to that voice since there have been lots of advancements in voice synthesis since he originally got this and yet he's chosen to keep this one.

I wonder if he'd accept the same voice, but with a different accent?

Should be just as possible in the present or near future.

Re: Lyrebird – An API to copy the voice of anyone

#252
post #101

Combined with Face2Face[1] live video impersonation, it is truly time to be very careful verifying videos or even live streams. https://www.youtube.com/watch?v=ohmajJTcpNk

To my knowledge, both of these particular projects are still a ways away from being used in any practical sense, let alone succeed at deceiving anyone.

You are right that we'll have to worry about this soon though. Likewise, verifying the identify of people we think we're talking to over video calls for example.

Re: Lyrebird – An API to copy the voice of anyone

#253
post #15

Earlier quoted context omitted.

They claim they only needed one minute of sample audio. We really need to start requiring that all public announcements (news, press releases, etc) are digitally signed and put into the blockchain.

But who would sign it?

Trump. We'd need to teach people to assume anything unsigned is fake.

Re: Lyrebird – An API to copy the voice of anyone

#254
post #183

Earlier quoted context omitted.

Woah that's kinda scary. What could we do to determine if a video is legitimate or not?

Mainly, practice critical thinking. Don't take anything at face value until it has been reconfirmed from many sources. At least that's what I do.

This is going to be increasingly key. And even then, it will be very difficult.

Books like "Trust Me, I'm Lying" reveal the lengths at which deception can occur. Though this book discusses deception that starts at the textual level (e.g. blogs), it is inevitable that these tactics will be translated to the video level once the technology catches up.

Also, "at face value" - Ha! ;)

Re: Lyrebird – An API to copy the voice of anyone

#255
post #74

Earlier quoted context omitted.

It doesn't sound too different from a voice coming over a walkie talkie or some kind of intercom. The problem might be that high frequencies, especially overtones, aren't properly constructed, but I'm certain that can be improved.

The main problem is that the algorithms don't yet know what to stress in a sentence. The problem is semantic, and not so much about the sound of the voice itself. You can synthesize someone's voice perfectly, but if it's stressing words incorrectly or not at all, it's not going to fool anyone. Then again, that's probably easier to work around by having humans annotate the sentences to be read.

pragmatic*, placing stress is less a problem of word meaning than it is of speaker adaptation for listener comprehension, emphasis, and prosodic tendencies.

Even then, I don't believe the issue is with stress. I believe that the voices sound robotic because they are using, and also admitting because it makes their results impressive in some sense, very few samples, "less than a minute" they claim. Triphones are usually what speech systems are trained on. The amount of triphones (3-phoneme-grams) to cover a language's phonemic inventory is huge (50 phonemes = 50! triphones, which could mean a few hours of audio, although many will not occur within the language given the phonotactics of the language).

Re: Lyrebird – An API to copy the voice of anyone

#256
post #6

This is pretty cool (although, I have no idea what other technologies exist for this kind of thing), but it's definitely not convincing enough to a human listener. This sounds like it might be convincing enough for some programs like "Hey, Siri" but it's not gonna convince your mom. You can listen to the samples on the page linked here and you can immediately tell that Obama and Trump don't sound quite human.

Well, the question is, do they just need to throw more computational power / training at this algorithm or is that the peak of their implementation? This is something Google has been working a lot on [1] and Baidu also recently posted about their results too [2]. We're definitely pretty close to passing the human detectable level. [1] https://deepmind.com/blog/wavenet-generative-model-raw-audio... [2] http://research…

I believe what makes the voices robotic is due to the little amount of audio they need to generate a "usuable" voice from the system.

Speech models usually use triphones, which turns out to be a huge amount of audio. This is particularly impressive because of how little data they need.

Google used their own datasets, which are most likely massive.

Re: Lyrebird – An API to copy the voice of anyone

#257
post #39

Earlier quoted context omitted.

I don't think you can claim that your game was voiced by either of them, but I don't see how using this would be any more infringing than using a tuned synthesizer.

He probably won't be able to claim the game is "voiced by", but maybe he can get away with saying it features "voices of"?

Actually i wasn't planning on saying either, just having the calm voice of Morgan Freeman tell me i need shoot the zombies or whatever the game will be, and then David explaining with fascination how such a poor shot could have survived this far whenever you miss

Re: Lyrebird – An API to copy the voice of anyone

#258
Curious choice to name a company & product with a name that sounds like "Liar Bird" when spoken. To me, that looks like they're fully embracing the concept that this can be used for nefarious purposes. If one of their goals is to bring attention that this technology exists and can be misused, the name reinforces that.

Re: Lyrebird – An API to copy the voice of anyone

#259
post #258

Curious choice to name a company & product with a name that sounds like "Liar Bird" when spoken. To me, that looks like they're fully embracing the concept that this can be used for nefarious purposes. If one of their goals is to bring attention that this technology exists and can be misused, the name reinforces that.

I don't see why audio manipulation would be any more nefarious than photo manipulation or video manipulation.

Plus, it doesn't even matter. You can write an article with fake quotes and people will believe it without even caring if there is an accompanying sound byte or not.

Re: Lyrebird – An API to copy the voice of anyone

#260
post #101

Combined with Face2Face[1] live video impersonation, it is truly time to be very careful verifying videos or even live streams. https://www.youtube.com/watch?v=ohmajJTcpNk

Without a doubt, our concept of personal identity will be completely unreliable within a few generations. Forget about privacy--we will soon have literally no way to verify who we're talking to.

Crypto would still work, and this tech isn't going to work face-to-face.
Post reply on HN