Live data from Hacker News

It’s Game over on Vocal Deepfakes

daringfireball.net

111–120 of 122 posts

Re: It’s Game over on Vocal Deepfakes

#112
post #34

Earlier quoted context omitted.

For some folks , they would just want to have a conversation with your favorite actor, author, celebrity psychologist just for fun.

Now there's an idea! Chat with a psychologist using your dead parents voices. I wonder what Freud would make of that.

No no no. Much better: read & grok all those unread books on my bookshelf, and then have looong conversations with me about their contents (and closely-related topics) at the times of my choosing. And when you quote from one of those books, don't be afraid to try to adopt the author's own voice.

Re: It’s Game over on Vocal Deepfakes

#113
post #103
post #62

Earlier quoted context omitted.

MacLuhan and others have already talked about this, from way back in the radio and TV generation. I don't say this to mean, "oh, and they were wrong because everything turned out OK" If anything, we have had generations that grew up with this dysfunction in a normalized way, and it took another advanced in tech to really see it again. (To take it further back into history, I think I now get what Cervantes was trying…

If anything TV and radio presented a larger shared reality, even if it was one filled with falsehoods. What is culture but a shared set of values? What we are growing now is a system that can give each and every one of us our own reality, and if it succeeds, what of society?

There are some fascinating areas to explore those questions from the perspective of shamans, mystics, and yogis.

Re: It’s Game over on Vocal Deepfakes

#114

Does anyone have suggestions on verbal ways to authenticate family members, for example, on the phone?

One way: Everyone has a code-word which they only share with humans, over "private" channels like direct voice calls. For example, I'm "pineapple", and you're "orange", when we're connected on a call, we share our code-words in order to "authenticate" the call. Another way: have a passphrase with a group of people that you share with each other, that is hard to guess but memorable. "The sky was green this morning" fo…

> Another way: start everything with a video call where they have to pass their hand in front of their face before switching to audio only.

and also they have to turn to profile.

Re: It’s Game over on Vocal Deepfakes

#115

Earlier quoted context omitted.

Maybe a use case for NFTs hah.

It's a use case for merkle trees. Since the mere existence of the hash at a reasonably well-attested time is sufficient, you don't need the rest. A handful of central roots of various trees is fine, morally equivalent to announcing things with an ad in the newspaper.

Now that I think about it, it seems to be a great idea - a verifiable record that some digital artifacts have existed at certain times. Of course it wouldn't be able to "prove ownership" or "prove being original" or other stupid ideas tokenbros throw around, that's impossible. But it wouldn't be profitable for the tokenbros and actually all the distributed attribute is not a requirement. We can probably set up a centralised append only DB which will periodically publish it's own hashes to attest that it wasn't tampered. Probably a useful tool for our deep fake future, the only question is how to scale it, who can be trusted and relied upon to input hashes in the system.

Re: It’s Game over on Vocal Deepfakes

#116
post #4

Once the generation times get to real-time, I dunno what's going to happen. I follow the voice acting community a lot and this is a big existential threat level of worry in that community. But I also see some positives for the narrative voice field, but at the expense of actual actors. The latest sequel to a favorite audio book series has the professional narrator pronouncing different character names and town names…

it’s already here. People on the chans are using a alexjones model to sing anime songs and it worked really well.

Alex Jones has been made to sing anime songs for over a decade: https://www.youtube.com/watch?v=ODZE5peUfWQ

Re: It’s Game over on Vocal Deepfakes

#117
post #70
post #47

Earlier quoted context omitted.

The audio directly in the post was for sure, but did you listen to the audio in the linked twitter thread? The only thing that gave me pause was the gpt text itself not the voice audio.

I had not listened to the audio in the twitter thread. It's a lot better and the only issue I notice is the pauses at the end of sentences don't seem quite right. I don't know if I'd notice the pauses though if I were not going into this with a suspicious mindset. After listening to the twitter thread I'm not convinced I could detect the difference if it was two random clips one generated and one not.

The Twitter audio had a distinct "key note address" cadence to me rather than being conversational. Of course probably most training voice came from such addresses.

Re: It’s Game over on Vocal Deepfakes

#118
post #115

Earlier quoted context omitted.

It's a use case for merkle trees. Since the mere existence of the hash at a reasonably well-attested time is sufficient, you don't need the rest. A handful of central roots of various trees is fine, morally equivalent to announcing things with an ad in the newspaper.

Now that I think about it, it seems to be a great idea - a verifiable record that some digital artifacts have existed at certain times. Of course it wouldn't be able to "prove ownership" or "prove being original" or other stupid ideas tokenbros throw around, that's impossible. But it wouldn't be profitable for the tokenbros and actually all the distributed attribute is not a requirement. We can probably set up a cent…

The idea predates bitcoin IIRC.

Re: It’s Game over on Vocal Deepfakes

#119
post #6

I tried ElevenLabs and it is truly amazing. You don't need much audio to train it and the final result can be incredible. I shared some of the snippets with friends and relatives and after the intitial scare (same as John) we agreed that the outcome wouldn't be different than a few years ago... We had impersonators for decades, so what have stopped political parties to hire an impersonator to create fake audios? The…

How does ElevenLabs perform for dialogues with emotions? Like for movies, TV shows and games? My impression about these voice generators was they're only good for Youtube tutorials and such. It was 2 years ago tho, I wonder how things changed.

What I have read people do is that they prime the text with something obviously emotional before hand, including punctuation, before their desired output. They then trim the audio they want to remove the prime phrase.

So something along the lines of:

"Remove this audio because it makes me super angry! Very angry! Now I say what I want to keep!"

Just having the AI read a script is often not enough alone without post processing and manipulation, often it is somewhat flat

Re: It’s Game over on Vocal Deepfakes

#120
post #31

I think he's right about this—things are going to get weird and ugly and people aren't prepared for what's coming. Here's my suggestion for safeguarding your sanity in the years ahead: find the people and the blogs you love, embrace RSS, block or avoid everything else, read old books, stop listening to podcasts, revert to email as the primary channel for occasional "social" correspondence, abandon all side projects t…

…go outside. This is the way. Hmm. Thinking in this context… Wearing a helmet even in a room of friends looks like overkill from anti-surveillance perspective, but could be a good protection from deep fakes: someone could still record your face with a hidden camera, but it will be useless for fakes because no one knows it's you and on the other hand everyone knows you don't walk around bare face.

[deleted]
Post reply on HN