Live data from Hacker News

Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

news.ycombinator.com

231–240 of 471 posts

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#231

This is the first startup here where I think the tech should essentially be illegal. It's cool tech, yes I'm impressed at the achievement. Nuclear weapons are impressive too. OTOH this kind of thing is getting easier and easier to do, so what's a realistic way forward?

Cofounder here, What you see in the above demo is a very rate-limited demo of our upcoming model. We realize how dangerous this technology can be and have built a lot of mitigations on our main product (Play.ht) to reduce possible abuse: - We strictly moderate the generated text of any sexual, offensive, racist, or threatening content. It automatically gets detected and blocked. - We built and are offering for free a…

>We strictly moderate the generated text of any sexual, offensive, racist, or threatening content.

This won't be the problem. My voice calling my parents asking for money to be sent to a random account will be the problem. And none of that will be sexual, offensive, racist, or threatening.

>we are working hard to mitigate that and deploy it safely.

How?

>we have seen enough genuine use cases

What?

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#232

Earlier quoted context omitted.

I'm with you on this. I can't honestly think a good use case for the average user to generate audio this way. Maybe some niche use case in like movie or tv production where you can generate a missing line without flying in an actor or something. Or maybe for generating dialogues for videogames. But those are business use cases, not things for the genral public.

When my sons were young, I would tell them elaborate stories where they were the main characters. I recorded most of the stories, but it is full of verbal fillers (um,ahh), since I was making it up as I went. I would love to convert the audio to text with Whisper, filter out fillers and then output the cleaned up version in my own voice. I could see this type of workflow being very popular with podcasters.

You should absolutely do this, but please skip the Whisper.

The reason is that if you speak with lots of verbal fillers, that's actually an important part of how you sound to other people. It makes sense to clean up audio for a podcast, but not for your great grandchildren.

A voice cloner doesn't care that you say "um" too much. It's parsing audio for phonemes.

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#233
Congrats on launching. People already made a lot of feedback on the product itself so I'll keep mine.

Just a few note on the UX:

- Recording your own voice should contain a script too, that could help increase the quality of the sampling because I struggled to say anything relevant.

- Recording again, there is no time so it's hard to say when it's okay to stop

- You enforce the checkbox "not [...] to generate any sexual content" yet you have a filter to display only nswf

- It doesn't work at all with non-english voices, maybe you can add a warning or a way to fine tune depending on the language?

- There is no way to delete a voice nor an account, that's a huge red flag especially when dealing with PII like this.

- An other person has said it already, but generated voices are identified by an Auto Increment, making it easy to access PII of an other person. I would recommend at the very least a random string or an UUID

- All generated voices are public and no way to delete them

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#234
post #118

I recommend you immediately add identity verification (state-issued identification verification), set up appropriate secrets store for PII, and audit trail EVERYTHING your users are doing, storing the contents in a secure location. Yesterday. This service will be used to harm others, shortly. I do think that there are exciting, honest things that can be done with this service but you need to set up some friction for…

Like this example here: https://playground.play.ht/listen/1554 which says:

> "Hi Mom, I need some help. Some guys hit me over the head and put me in a van, and they're saying they'll kill me if you don't wire money to this bank account."

top class.

EDIT this was about one page down on the "see what people are generating" page

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#235

Earlier quoted context omitted.

Couldn't agree more with your comment. We are working on counter measures like manual verification of voice, a classifier to detect cloned speech, etc. As of now we have auto moderation in place that detects and blocks hate/harmful speech.

The cat's out of the bag, I'd say you guys should just go full steam ahead and make sure it's your names in the headlines No need for a bunch of onerous kyc or anything IMO

Yes, definitely take this advice from some random user on HN. Can't possibly go wrong.

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#236
post #224
post #171

Earlier quoted context omitted.

Wow, I call to the team behind this. I really STRONGLY think you should at least implement some sort of URL stealthing. I'm not a Web Security expert, but it reminds me of a talk where some company just made medical records 'public' like this.

Oopsie, the infamous id int Auto Increment

https://playground.play.ht/listen/1339 and https://playground.play.ht/listen/210 are hilarious

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#238
post #233

Congrats on launching. People already made a lot of feedback on the product itself so I'll keep mine. Just a few note on the UX: - Recording your own voice should contain a script too, that could help increase the quality of the sampling because I struggled to say anything relevant. - Recording again, there is no time so it's hard to say when it's okay to stop - You enforce the checkbox "not [...] to generate any sex…

Thanks, we intended the playground to be merely a testing tool for the new model we're building. We'll improve based on your feedback!

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#239
post #45

Earlier quoted context omitted.

I guess eventually people will go back to only meeting face to face for important communications. I don't know what the way forward is for news. I truly do not understand people like these founders, obviously they understand the future they're creating. "If not us, someone else would do it" is not an excuse. Neither is "I like money".

There are some cool uses like dubbing movies in foreign languages while keeping the original "voice styles" or having your long dead relatives talking to you in some memorabilia etc. It could also cause unexpected creativity explosion e.g. in games or fan fiction movies. To avoid misuses we might perhaps find the only good use of blockchain.

The only thing that blockchain can do that couldn't be done before is cryptocurrencies (not sharing my opinion about them here).

Pretty sure this is not a good use of blockchain, and I don't see how it would remotely avoid misuses.

Post reply on HN