Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio
451–460 of 471 posts
Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio
#452Congrats on launching. People already made a lot of feedback on the product itself so I'll keep mine. Just a few note on the UX: - Recording your own voice should contain a script too, that could help increase the quality of the sampling because I struggled to say anything relevant. - Recording again, there is no time so it's hard to say when it's okay to stop - You enforce the checkbox "not [...] to generate any sex…
The terms of service is terrifying for anybody who has a voice or anything of value they want made into speech > you automatically grant, and you represent and warrant that you have the right to grant, to us an unrestricted, unlimited, irrevocable, perpetual, non-exclusive, transferable, royalty-free, fully-paid, worldwide right, and license to host, use, copy, reproduce, disclose, sell, resell, publish, broadcast, r…
Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio
#453Earlier quoted context omitted.
Two thousand eight = American English Two thousand and eight = UK and Australian English
I'm based in Europe and a native English speaker, I thought I was aware of most of the differences between UK/US English. I can't believe I have worked with Americans for decades and never noticed this. Live and learn!
Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio
#454Earlier quoted context omitted.
I think this in an area where there are many more malicious use cases than legitimate. It's like spyware developers that claim their software is for remote administration of computers you own.
Don’t take this personally, but I don’t think you’ve thought too hard about what you can do with this technology. For instance, you could create audiobooks for every book ever published. You can change scripts for movies after shooting has already happened. Indie game developers can now afford high quality voices in their games. Even AAA games like The Elder Scrolls can vastly expand their in-game voice variety. I th…
Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio
#455Earlier quoted context omitted.
I'm with you on this. I can't honestly think a good use case for the average user to generate audio this way. Maybe some niche use case in like movie or tv production where you can generate a missing line without flying in an actor or something. Or maybe for generating dialogues for videogames. But those are business use cases, not things for the genral public.
When my sons were young, I would tell them elaborate stories where they were the main characters. I recorded most of the stories, but it is full of verbal fillers (um,ahh), since I was making it up as I went. I would love to convert the audio to text with Whisper, filter out fillers and then output the cleaned up version in my own voice. I could see this type of workflow being very popular with podcasters.
You can use Descript, CleanVoice, and other tools to achieve exactly what you just said, in a few minutes, from just the original recording.
Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio
#456Earlier quoted context omitted.
But what is a realistic way forward? Do you think that scammers won't have this technology in 2 years? Can we really prevent any illegal use of neural networks at this point? With weapons that you actually have to physically buy, you can intervene on a country level (to some degree). But already with those 3D printed ones, we are basically doomed. Of course it's a tragedy of the commons type of situation. But banning…
Of course it's a tragedy of the commons type of situation TOTC is about resource depletion. GYI. It's not applicable here.
Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio
#457Earlier quoted context omitted.
Pretty sure that's something that ought to have been discussed before any of this ever started, but you know, scientists, could, should. I look forward to the chaos and destruction and all these "brilliant" software developers wringing their hands saying they couldn't possibly have imagined such horrible outcomes from their fun money-making venture that just so happened to undermine the concept of a shared reality.
I think we all thought we'd be able to come to those decisions on a more gradual timeline. The breakneck pace of AI breakthroughs over the past few years have revealed: not so much.
All's well that ends well, though. We simply don't have the resources to continue this "breakneck pace".
Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio
#458Earlier quoted context omitted.
Don’t take this personally, but I don’t think you’ve thought too hard about what you can do with this technology. For instance, you could create audiobooks for every book ever published. You can change scripts for movies after shooting has already happened. Indie game developers can now afford high quality voices in their games. Even AAA games like The Elder Scrolls can vastly expand their in-game voice variety. I th…
You're not wrong, all of that is great. But the capacity for even more spam calls and scammers generating fraudulent content with cloned voice samples will be an immensely annoying issue. Anyone who owns a phone in the modern age absolutely cannot be a Pollyanna when looking at this technology. There are real issues that must be addressed.
Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio
#459Earlier quoted context omitted.
Yes, we are working on making the API pay as you go soon. Thanks for the feedback!
Another note: the share view on the clips doesn't include any way to get the actual link to the file. I imagine most people want the actual link so they can have more control over how and where they share it.
Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio
#460It was creepy af considering that the recording I used had none of those elements (no background noise at all).