Live data from Hacker News

Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

news.ycombinator.com

71–80 of 471 posts

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#71

https://play.ht/app/voice-cloning > Clone a voice now Pops a modal: Try Voice Cloning for Free! Enter a credit card for $0.00/mo with no other information on screen Bounce. Why not let me play around with it a little without asking for a credit card?

I think if you are cloning voices, you should be required to have a credit card or some other KYC identifier. Even if it's free. This kind of highly abusable tech should have a paper trail IMO.

yeah because that's working great for crypto lol

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#73
post #50

Earlier quoted context omitted.

Making then the illegal would accomplish nothing since it's already out in the wild. You can generate audio with high quality on fine-tuned versions of Tortoise TTS, which was originally trained on a cluster of NVIDIA 3090's, so it's within reach for any smart person to train a from-scratch model on consumer hardware. Realistically? We have to accept that this tech exists and there will be both positive and negative…

> Making then the illegal would accomplish nothing since it's already out in the wild. Not true. Making it illegal wouldn't make it nonexistent -- that's true. But making it illegal would provide at least some method of mitigating some of the harm. That's more than what we have right now. > We have to accept that this tech exists and there will be both positive and negative outcomes from it. Of course. But that doesn…

> Not true. Making it illegal wouldn't make it nonexistent -- that's true. But making it illegal would provide at least some method of mitigating some of the harm.

It's already illegal to impersonate someone to steal money or scam them, and those laws were on the books before computers existed.

> Of course. But that doesn't mean it's futile to try to reduce the negative outcomes.

You can run something on a consumer GPU and it's every bit as good if you know how to dial it in. By the end of the year you'll be able to download a nicely packaged "voice cloner" from a torrent that runs on a cheap laptop. IMHO any effort on regulation is far better spent informing people rather than trying to put the cat back in the bag.

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#74

Earlier quoted context omitted.

>Introducing the National Postal Service - send a letter to anyone for a nominal fee. No need for a personal courier, armed escort, or patrician status. >This is the kind of thing that should be illegal. Now, any Plebian could essentially write a letter to anyone, impersonating anyone. Forged letters could drag us into a war with Persia - for Jupiter's sake!

This is a hilariously bad attempt at discrediting the original argument. There's a vast difference between forging a letter and replicating the unique vocal fingerprint of any human being, on demand. I suppose if we approach the point that we can create robotic clones of anyone, anywhere, that look, sound, and move like anyone on the planet, that will be just like the post office too, right?

What are some of the differences? Besides the glaringly obvious text vs audio. I mean prior to telegraphs if I got a letter from my sweet heart with a lock of hair or something and a request for funds I’d probably believe it, especially if it took days or weeks to communicate back and forth?

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#75

How is the latency for real-time TTS? I remember kicking the tires several months back but went with one of the big 3 cloud providers since they had lower latency. I also like that the cloud provider supports SSML and I can explicitly configure the emotion, whereas Playht dynamically changed the emotion based on context of the text.

The latency is not real-time yet but we're working on getting it to near real time. Regarding controlling the voice, we've added a few params like rate, voice guidance and temperature but for the most part the emotion is dependent on the text for now.

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#76
post #31
post #4

How do you assert that the cloned voice has been truly permitted by the voice owner? I've had my voice cloned without my consent by other people using Descript and Eleven Labs. What is your process for verifying consent?

When I tried this service previously, you had to read (out loud) something saying that you were giving consent.

[deleted]

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#79
post #63
post #31

Earlier quoted context omitted.

When I tried this service previously, you had to read (out loud) something saying that you were giving consent.

I'd be curious what the false positive rate on that is. Can you clone anyone's by collecting a set of ten voices with unique timbre reading the required statement plus pitch control to get close enough? A hundred? Or can you trick the neural net by giving it something that sounds like white noise to humans until the NN triggers in the right way and goes "ok yep that's a match, you're authorised now"? Probably not som…

Yeah, good point - don't know. When I tried I actually did get a (personal?) email saying that it didn't match closely enough. After uploading another sample (based on a different text) it went through.

I like your idea of just training on the consent text! That wasn't the case when I tried it as you needed around 3h (optimally) of training data.

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#80

This is the first startup here where I think the tech should essentially be illegal. It's cool tech, yes I'm impressed at the achievement. Nuclear weapons are impressive too. OTOH this kind of thing is getting easier and easier to do, so what's a realistic way forward?

I'm with you on this. I can't honestly think a good use case for the average user to generate audio this way. Maybe some niche use case in like movie or tv production where you can generate a missing line without flying in an actor or something. Or maybe for generating dialogues for videogames. But those are business use cases, not things for the genral public.

Pretty much any hobby video game development or animation?
Post reply on HN