Live data from Hacker News

Lyrebird – An API to copy the voice of anyone

lyrebird.ai

11–20 of 311 posts

Re: Lyrebird – An API to copy the voice of anyone

#12

This site has a "demo" section featuring only Soundcloud clips. Uses to much the present tense "In a world first, Montreal-based startup Lyrebird today unveiled" and "Record 1 minute [...] and Lyrebird can [..]Use this key to generate anything" but has no actual product or beta version. Adobe had a much more impressive sneak peek of a similar product called VoCo: https://www.youtube.com/watch?v=I3l4XLZ59iw

To be fair, that demo could have been staged whereas we can be pretty darn sure those aren't Trump's actual words.

Re: Lyrebird – An API to copy the voice of anyone

#13
post #6

This is pretty cool (although, I have no idea what other technologies exist for this kind of thing), but it's definitely not convincing enough to a human listener. This sounds like it might be convincing enough for some programs like "Hey, Siri" but it's not gonna convince your mom. You can listen to the samples on the page linked here and you can immediately tell that Obama and Trump don't sound quite human.

Text to speech is still pretty distinguishable as not-human, and that seems like an easier problem (only has to work for one specific voice, not an arbitrary voice). So just on the basis alone I wonder if this isn't still a ways out

Re: Lyrebird – An API to copy the voice of anyone

#15

This is how a lot of tech companies make proper text2speech, this was just done using the vast amount of audio that's out there for these people. Soon Trump will use this to state that things he's said are fake news. God help us all.

They claim they only needed one minute of sample audio.

We really need to start requiring that all public announcements (news, press releases, etc) are digitally signed and put into the blockchain.

Re: Lyrebird – An API to copy the voice of anyone

#18
post #6

This is pretty cool (although, I have no idea what other technologies exist for this kind of thing), but it's definitely not convincing enough to a human listener. This sounds like it might be convincing enough for some programs like "Hey, Siri" but it's not gonna convince your mom. You can listen to the samples on the page linked here and you can immediately tell that Obama and Trump don't sound quite human.

Well, the question is, do they just need to throw more computational power / training at this algorithm or is that the peak of their implementation?

This is something Google has been working a lot on [1] and Baidu also recently posted about their results too [2]. We're definitely pretty close to passing the human detectable level.

[1] https://deepmind.com/blog/wavenet-generative-model-raw-audio...

[2] http://research.baidu.com/deep-voice-production-quality-text...

Post reply on HN