Live data from Hacker News

OpenVoice: Versatile Instant Voice Cloning

arxiv.org

71–80 of 207 posts

Re: OpenVoice: Versatile Instant Voice Cloning

#71
post #51

And suddenly it becomes a bit weird: https://docs.myshell.ai/tokenomics Tokenomics Disclaimer: MyShell is currently in the testing phase, and the content of the whitepaper may be subject to change in the future. $SHELL is the token used for user incentive, governance and in-app utility. The total supply of $SHELL is 1,000,000,000

[deleted]

Re: OpenVoice: Versatile Instant Voice Cloning

#73
post #52

Earlier quoted context omitted.

Unifying a voice in tutorial videos so that the difference in voice does not distract the learner. Auto non-toxic rephrasing of online chat in video games, let people hear their voice but paraphrase what they said in a manner that doesn't turn the platform into a cesspit. Cloning your own voice so that you can turn a script into audio without 50 takes and then having to remove a million Ums and errs.

> Auto non-toxic rephrasing of online chat in video games, let people hear their voice but paraphrase what they said in a manner that doesn't turn the platform into a cesspit. that feels very orwellian

George Orwell — 'If you want a picture of the future, imagine a boot stamping on a human face—for ever.'

I think this is closer to the direction of Huxley in Brave New World, where a deeper understanding of how to manipulate without brute force creates a very different dystopian society than 1984.

Re: OpenVoice: Versatile Instant Voice Cloning

#75
I commend the authors on making this easy to try! However it doesn't work very well for me for general voice cloning. I read the first paragraph of the wikipedia page on books and had it generate the next sentence. It's obviously computer generated to my ear.

Audio sample: https://storage.googleapis.com/dalle-party/sample.mp3

Cloned voice (converted to mp3): https://storage.googleapis.com/dalle-party/output_en_default...

All I did was install the packages with pip and then run "demo_part1.ipynb" with my audio sample plugged in. Ran almost instantly on my laptop 3070 Ti / 8GB. (Also, I admit to not reading the paper, I just ran the code)

Re: OpenVoice: Versatile Instant Voice Cloning

#76
post #20

I had to phone my bank, which is one of the bigger players in the UK high street market, a couple of days ago. They're still encouraging me to enroll in their idiotic "my voice is my password" programme. At this stage in the evolution of AI, that feels simply negligent.

Fidelity Investments just did something even worse ~a week ago - It asked me to reply to a few questions, then announced that I'd just been enrolled in it's voice identification program (or whatever they call it). Now I've got Just Another Item on my ToDo list, to get that undone. Gawd, does every company promote it's stupidest people to management?

I don't know if GDPR (or any of its cousins) applies to you, but this kind of thing sounds exactly like the sort of thing it's supposed to outlaw.

Re: OpenVoice: Versatile Instant Voice Cloning

#77

Can someone give me a practical use case where this adds a net benefit to society?

No. The real answer is yes, I could probably come up with some contrived examples, like I lost my voice in a freak LLM accident and now want to clone my old voice. But this doesn't (you don't?) really need a net benefit reason to figure it out and publish it. Because why? I assume, because "this shouldn't exist!" which is just a more palatable wa to phrase "won't someone think of the children". Society doesn't benefi…

> Why does it need a practical reason?

To at least give us something as a consolation for all the havoc all sorts of deep fakes will wreak on societies. It's like asking what a knife can be used for other than murder. It's a valid question.

Re: OpenVoice: Versatile Instant Voice Cloning

#78
post #36

Earlier quoted context omitted.

Elevenlabs has been around for a while now. Genie has been out of the bottle for a bit, and the sooner the notion that anything digital can be easily faked seeps into the wider consciousness the better. Trust nothing.

I've seen some prank calls (a YouTuber cloned Tucker Carlson's voice and called Alex Jones) but he just had a sound bank with a few pre-generated lines and it fell apart pretty quickly. At least for now there's too much lag to do a real time conversation with a cloned voice. Speech to Text > LLM Response > Generate Audio If that time can shrink to subsecond, I think there'll be madness. (Specifically thinking of roma…

At last summer's WeAreDevelopers World Congress in Berlin, one of the talks I went to was by someone who did this with their own voice, to better respond to (really long?) WhatsApp messages they kept getting.

It worked a bit too well, as it could parse the sound file and generate a complete response faster than real-time, leading people to ask if he'd actually listened to the messages they sent him.

Also they had trouble believing him when he told them how he'd done it.

Post reply on HN