Live data from Hacker News

Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

news.ycombinator.com

401–410 of 471 posts

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#401
post #53

I'm having a hard time coming up with a non-nefarious use case for this.

Anything written can be listened to with this tech. Any news article, any short story, a draft of a piece of writing you're working on. There is too much text for human beings to read it all.

And all AI bots are here to generate even more text. :( We will need to rethink and reevaluate lots of things that we are used to.

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#402
Hey Mahmoud and Hammad

We love play.ht and we're already using it in our new start up called Aloudable. We convert email newsletters into podcasts (for now).

This is our MVP if you'd like to sign up: https://aloudable-frontend.vercel.app/

Will

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#404

Hey HN, we are Mahmoud and Hammad Are you though? You might just be computer-generated. While I'm very impressed with this technically (and as a pro-audio person I feel validated to see my predictions of a few years back coming true so dramatically), I don't see anything about risk management in here. Your tech absolutely will get used by scammers, given the overabundance of voice data on the open internet. How are y…

Wow, I haven't even thought of that. Imagine this being used together with a chatgpt equivalent. Scam rates are going to go through the roof.

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#407
post #356

Given the very ( very , VERY) obvious concerns associated with malicious deployment of this tech, and the minimal/largely ineffectual countermeasures deployed by the founders, what surprises me the most is that YC gave this startup its stamp of approval. It used to be that they offered at least a basic sanity check to anything they funded. Is this now getting lost as they scale up their funding operations?

The constant worry about malicious deployment is so tired in my opinion. The technology to clone voices exists. Your trust is audio recording should already be shaken. Trying to hobble this product on the grounds of "it's dangerous" just serves to limit creativity.

asking someone to take basic precautions like “have a fire extinguisher” or “hazmat labeling” just gets in the way of innovation!

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#408
post #53

I'm having a hard time coming up with a non-nefarious use case for this.

I'm using this kind of technology for temporary voice tracks in animated shorts. I'd really like something like Img2Img for voices so I can translate a performance to an arbitrary (synthetic) voice.

Tortoise TTS can do this. You just pass it your example as a conditioning latent.

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#409
post #118

I recommend you immediately add identity verification (state-issued identification verification), set up appropriate secrets store for PII, and audit trail EVERYTHING your users are doing, storing the contents in a secure location. Yesterday. This service will be used to harm others, shortly. I do think that there are exciting, honest things that can be done with this service but you need to set up some friction for…

How is it that > I recommend you immediately add identity verification (state-issued identification verification) and > The genie is already out of the bottle. The degree of effort to put this together is low enough that it will be replicated around the world. are thoughts that end up in the same post? If the genie is out of the bottle, it’s your proposed solution that everybody that runs a model like this implements…

It's more of a "Cover Your Ass with Paper" type thing.

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#410
post #53

I'm having a hard time coming up with a non-nefarious use case for this.

We have been seeing some of these genuine use cases: youtube creators, audiobooks, elearning videos, podcasts, commercials, dubbing, and gaming.

No-one is going to listen to an audiobook made with this. It's still fundamentally just TTS.
Post reply on HN