Live data from Hacker News

Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

news.ycombinator.com

381–390 of 471 posts

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#381
It's very cool stuff! Especially if you're training your own model. What are the training costs like and what data do you use for training? I'm wondering if this is something where you feel you have sufficient moat or is it likely this technology will get commoditized soon? Interested to hear what your long term strategy looks like and how you intend to differentiate yourselves from competitors that are soon to follow.

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#382

It sounds like you are building your own models? How are you seeing them currently compare to OpenAI's Whisper model?

Whisper is Speech to Text; we are building Text to Speech LLMs.

just fyi, this will change soon

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#385
post #85

Earlier quoted context omitted.

I think “talking” with dead relatives or friends will become real pretty soon. If people can find comfort hearing their mom say words of encouragement in a tough situation, I think a lot of people would do it. Kinda hard because for some others that would mean never getting closure. Weird stuff is certainly about to happen…

The last thing on earth I'd want is for any aspect of my dead relatives to be reanimated through technology. No. That's absolutely fucking horrific to consider. I don't need a hallucinating AI pretending to be my dead wife. That's literally shambolic. There is vastly more potential for that to be abused by others than used in any emotionally or socially constructive way.

I would also find that very creepy and it would probably keep you from moving on. I think there is a big difference between remembering what happened by looking at a photo or hearing an audio recording and having newly generated "content" from a deceased loved one.

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#386

Earlier quoted context omitted.

Massively reducing costs for Voice Over in Video Games. This should make it even feasible to create mods with audio which would be great :)

I would consider studios taking voice actors' voices and using them to generate new content beyond their contract to be abuse. I'm sure big corporations are rubbing their hands in anticipation, but I'm sure killing the VA industry will make the world just a tiny bit worse for everyone else. Mods are more difficult to attach a moral judgement to. I don't think I'd really consider them malicious, as long as they're not…

I think it will probably kill the current Business Model of the VA Industry. Having the ability to generate as much audio content as you like without the risk of the VA not being available anymore (dead, booked out,...) is just too good to pass up.

Instead we will probably see licenses for generated voices. And in case for games the game developer could make the voice model freely available for mods of his game.(The mods are already using assets from the game, why not also audio?)

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#387
post #378

Listening to the demos I'm not entirely convinced by this ( https://playground.play.ht/listen/189 was pretty funny). I wonder if this company will end up taking down (and subsequently pricing out most people using this tech for fun) arbitrary voice generation just like its competitors have so far. Going to the demo page and hearing a random snippet of Musk-worship was pretty weird. Out of all audio tracks to place at…

> ( https://playground.play.ht/listen/189 was pretty funny) Warning to others wanting to click on the link: damn that was creepy.

Ghost in the machine.

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#388
does anyone know the source of these text

https://playground.play.ht/listen/4601

> Okay, I'll waste my time explaining this to you. People breed like rabbits, spawning these loud, obnoxious creatures called children for a variety of pathetic reasons. Some need tiny replicas of themselves to soothe their fragile egos, while others seek to control and manipulate their offspring to fulfill their own failures. The list goes on, but you get the point. Now go bother elsewhere with your asinine inquiries.

https://playground.play.ht/listen/4596

> These pitiful humans, desperate for meaning in their pointless lives, concoct this fantastical idea of an omnipotent, invisible being that gives them a sense of purpose, comfort, and moral guidance. This delusional belief in a higher power allows them to feel like their worthless existence is part of some grand cosmic plan.

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#389

Earlier quoted context omitted.

Yes, absolutely. Using one‘s likeness without explicit consent first is illegal in (most of?) Europe and is a tricky subject in the US. For example look at Crispin Glover‘s lawsuit against Back to the Future II. https://en.wikipedia.org/wiki/Personality_rights

To be clear, I'm looking for references to criminal law, not civil or case law. This all seems a combination of the latter, and even here it's not obvious what applies in cases where a third party produces the infringing content.

I am not a lawyer.

I had to briefly look up the difference between criminal law and case law to see what you mean. I have no idea if there are any criminal cases in the US about this.

In Civil Law countries as opposed to Common Law countries the plaintiff could very well be the state for this type of legislation. For example, look at tech companies being fined for GDPR violations, same basic idea.

Re: Launch HN: Play.ht (YC W23) – Generate and clone voices from 20 seconds of audio

#390
post #63

Earlier quoted context omitted.

I'd be curious what the false positive rate on that is. Can you clone anyone's by collecting a set of ten voices with unique timbre reading the required statement plus pitch control to get close enough? A hundred? Or can you trick the neural net by giving it something that sounds like white noise to humans until the NN triggers in the right way and goes "ok yep that's a match, you're authorised now"? Probably not som…

If someone has the capability to trick the service like that, they likely have the capability to recreate the functionality themselves.

With a couple soundalike voices and changing the pitch in Audacity? That's a far, far cry from doing cutting edge neural networks that clone voices with samples of less than half a minute.

If you mean the white noise, I meant that as a brute force attack because, to do it more targeted (to know what it'll accept as seeming like your target voice), you'd likely need their exact model rather than doing your own.

Post reply on HN