Live data from Hacker News

Lyrebird – An API to copy the voice of anyone

lyrebird.ai

161–170 of 311 posts

Re: Lyrebird – An API to copy the voice of anyone

#161

Last week on BBC Radio 4 I heard of a woman who was losing her voice through disease (MND maybe?), a similar system was being anticipated and she was saving voice samples to seed it with. She had been a singer and strongly identified her self with her voice, she wanted to be able to use a speech synthesis system that had her own voice pattern. Apologies if this was already mentioned, but it seems to be a use others h…

This is actually quite inspiring!

Re: Lyrebird – An API to copy the voice of anyone

#162
post #101

Combined with Face2Face[1] live video impersonation, it is truly time to be very careful verifying videos or even live streams. https://www.youtube.com/watch?v=ohmajJTcpNk

just wait till _the daily show_ and _last week tonight_ get a hold of this!

Re: Lyrebird – An API to copy the voice of anyone

#163

Last week on BBC Radio 4 I heard of a woman who was losing her voice through disease (MND maybe?), a similar system was being anticipated and she was saving voice samples to seed it with. She had been a singer and strongly identified her self with her voice, she wanted to be able to use a speech synthesis system that had her own voice pattern. Apologies if this was already mentioned, but it seems to be a use others h…

I was just thinking that Stephen Hawking would perhaps be interested in using this to replace his current voice synthesizer (feeding in old interviews of him when he could talk). He has said that he has adopted the current voice since he has associated it with his own, but I wonder if he would prefer his old actual voice.

Re: Lyrebird – An API to copy the voice of anyone

#165
post #163

Last week on BBC Radio 4 I heard of a woman who was losing her voice through disease (MND maybe?), a similar system was being anticipated and she was saving voice samples to seed it with. She had been a singer and strongly identified her self with her voice, she wanted to be able to use a speech synthesis system that had her own voice pattern. Apologies if this was already mentioned, but it seems to be a use others h…

I was just thinking that Stephen Hawking would perhaps be interested in using this to replace his current voice synthesizer (feeding in old interviews of him when he could talk). He has said that he has adopted the current voice since he has associated it with his own, but I wonder if he would prefer his old actual voice.

I think Hawking is now so firmly tied to that voice that he would probably never switch for public speaking engagements and the like.

I could see him doing such a switch for personal interactions.

Re: Lyrebird – An API to copy the voice of anyone

#166
post #101

Combined with Face2Face[1] live video impersonation, it is truly time to be very careful verifying videos or even live streams. https://www.youtube.com/watch?v=ohmajJTcpNk

Woah, reminds me of Total Recall for some reason... looks like a special effect from the 80s when actual speaking occurs, but it's very close!

Re: Lyrebird – An API to copy the voice of anyone

#167

Cooler : http://www.dtic.upf.edu/~mblaauw/IS2017_NPSS/ https://arxiv.org/abs/1704.03809

This model is quite cool, but also quite a bit different than what lyrebird.ai is doing. NPSS has a lot of extra information in the control inputs about pronunciation and timing (the part-of-phoneme timer feature) - this means that most of the "hard parts" (in my opinion) for naturalness are control inputs to NPSS/WaveNet style models, rather than variables the model must generate globally and consistently as in lyrebird. At generation time NPSS appears to generate each component autoregressively as well, but I am not clear on whether the demo samples do this or if they use "true" values for f0 at least - what forces the model to sing the exact same melody, if many melodies are possible given the underlying audio information?

Also note that NPSS has some amount of post-processing, at least reverb and perhaps other common musical mixing - we don't really know how these samples are generated, and I have a hard time decyphering exactly what inputs are required, and what are generated from the paper alone. However, I really, really, really like NPSS - I just don't think the comparison you are making is valid here.

These features (f0, duration, pronunciation) are some of the most difficult things to learn to model from datasets of speech and text directly, and I am not sure how they got the subset used (I think only f0 and pronunciation/phoneme) for this NPSS model. Giving creators fine-grained control of the performance (as in NPSS) is quite cool, and if these systems can get fast enough I think the possibilities are really exciting. The same things could likely be done with lyrebird as well - there is no real "tech reason" you couldn't add more conditional inputs, with finer grained information/control.

The key part in my mind is deciding what amount of complexity to show to a user, and what amount to try and capture inside the model - some people may want to control (for example) duration and f0 directly for a performance, while others may want to just upload clips to an API and get reasonable results back, with less ability to control each sample (they can still curate themselves for the "best" samples). Lyrebird.ai is handling the latter case, while the former case would require quite a bit more intervention from the average user, almost becoming like an instrument ala the original voder [0]. However, you could potentially have both approaches as a kind of beginner/advanced mode, but advanced mode needs a user interface, and probably near-realtime feedback.

I used to really strongly believe that the audio model was going to be the hard part of "neural" TTS (blame my background in DSP perhaps), but post-WaveNet the game has really changed a lot - conditional audio models are something we are starting to know how to do pretty well.

The text pipeline of most TTS systems is still the craziest part in my mind, check out a "normal" feature extraction of 416 hand-specified features [1]! These extractions can be upwards of 1k features per timestep/frame, and generally require a lot of linguistic knowledge to specify for new languages. It seems (given Alex Graves' demo [2], char2wav [3], tacotron[4]) that we are making progress on learning this information directly from text, which in my mind is a key breakthrough for TTS in languages besides English, where lots of work on English pronunciation has been done already and is generally available.

[0] https://www.youtube.com/watch?v=TsdOej_nC1M

[1] https://github.com/CSTR-Edinburgh/merlin/blob/master/misc/qu...

[2] https://www.youtube.com/watch?v=-yX1SYeDHbg&t=38m00s

[3] http://josesotelo.com/speechsynthesis/

[4] https://google.github.io/tacotron/

Re: Lyrebird – An API to copy the voice of anyone

#168

As noted in other comments, all the samples still sound very robotic, so this is probably "just" a method to tune the parameters of an existing voice synthesizer to mimic a real persons voice as much as it allows.

That's exactly what it sounds like. The same old mediocre TTS with voices modified to mimic specific well-known voices.

It's impressive for what it is, but a lot of people here seem way too excited. This isn't any kind of breakthrough, and only the shortest hand-picked snippet would fool anyone.

Re: Lyrebird – An API to copy the voice of anyone

#169
post #101

Combined with Face2Face[1] live video impersonation, it is truly time to be very careful verifying videos or even live streams. https://www.youtube.com/watch?v=ohmajJTcpNk

just wait till _the daily show_ and _last week tonight_ get a hold of this!

..then they'll finally be able to play audio of republicans contradicting themselves! :p

Re: Lyrebird – An API to copy the voice of anyone

#170
post #163

Earlier quoted context omitted.

I was just thinking that Stephen Hawking would perhaps be interested in using this to replace his current voice synthesizer (feeding in old interviews of him when he could talk). He has said that he has adopted the current voice since he has associated it with his own, but I wonder if he would prefer his old actual voice.

I think Hawking is now so firmly tied to that voice that he would probably never switch for public speaking engagements and the like. I could see him doing such a switch for personal interactions.

Indeed, he doesn't like other voices and has always fallen back to using the same voice:

> "The voice I use is a very old hardware speech synthesizer made in 1986," he said. "I keep it because I have not heard a voice I like better and because I have identified with it."

http://usatoday30.usatoday.com/life/people/2006-06-15-hawkin...

Post reply on HN