Earlier quoted context omitted.
"Hey mom, it's X. Sorry to leave this voice mail, I got in a bit of a car accident and lost my phone. I need to get back home, could you send $$$$ to ####?"
Just have to establish real life security questions. They did this in one of the Harry Potter books (the sixth one?) since it is easy for them to impersonate other people. If you can’t tell me my favorite flavor of jam then you aren’t getting my money.
It’s Game over on Vocal Deepfakes
101–110 of 122 posts
Re: It’s Game over on Vocal Deepfakes
#102Earlier quoted context omitted.
The difference is the scale, speed, and cost. You have to at least find an impersonator who can do a passable impersonation of the target, which may be a tall order. They can't impersonate everyone perfectly, and the really good ones are probably busy and/or expensive. Anyone can use this. Kids could use it.
Is it? The examples given in the OP are remarkably uninspired, specifically mentioning "a recording of Joe Biden forgetting his own name or what year it is, or Kamala Harris claiming to be running an abortion clinic." At that level of course you can already easily find a high-quality impersonator if you wanted to. What kinds of stuff are kids going to use this for that is dangerous? Fake their parents to call in sick…
Re: It’s Game over on Vocal Deepfakes
#103Earlier quoted context omitted.
It's kind of amazing we're not seeing more of this already. It's a trivial application of the tech. We've already achieved a culture where the more batty something is the more likely people are to believe it, and now we have an infinite, personalized quackery generator. Without a massive cultural revolution putting the brakes on this, which I can't imagine coming from anywhere, I fear we're barreling into a future wh…
MacLuhan and others have already talked about this, from way back in the radio and TV generation. I don't say this to mean, "oh, and they were wrong because everything turned out OK" If anything, we have had generations that grew up with this dysfunction in a normalized way, and it took another advanced in tech to really see it again. (To take it further back into history, I think I now get what Cervantes was trying…
What we are growing now is a system that can give each and every one of us our own reality, and if it succeeds, what of society?
Re: It’s Game over on Vocal Deepfakes
#104Earlier quoted context omitted.
How does ElevenLabs perform for dialogues with emotions? Like for movies, TV shows and games? My impression about these voice generators was they're only good for Youtube tutorials and such. It was 2 years ago tho, I wonder how things changed.
I put some Shakespeare into the test on their main page and it was surprisingly good at times. Hearing an AI emote "Out, out, brief candle!" makes me feel strange. This has to replace voice actors in gaming at least. You could change the text as needed and perhaps just add some hints to get the emotion you want.
Re: It’s Game over on Vocal Deepfakes
#105So who's going to build the sex phone line where you can talk to celebrities like Marilyn Monroe?
The problems I see, in increasing order of difficulty are: smalltalk; memory; and dirty talk.
1st. A lot of the language models are very bad at pointlessness smalltalk that fills time, they aren’t very good at coming up with related conversational prompts or diversions to continue talking about without being carefully steered and purported by an extra layer of supporting software. The deliberately oversimplified example is that ChatGPT waits for you to say something, and will never just ask how your day was. This isn’t particularly difficult but it’s a case of it’s going to feel a lot like sophisticated NPC smalltalk in a video game for a while I suspect due to the need to drive it from a second system that monitors the conversation and provides the supporting timers to keep nudging along a conversation the way they do in video games
2nd. The majority of the models have limited memory and while we’re seeing clever tricks pop up all the time now about how to feed things back into the models prompts in order to keep it focused on a topic or to bring things back up later, it’s going to take another layer of software that try’s to identify key information and store that in a secondary system and then somehow contextually identify opportunities to mention it in order for us to have a chat bot that can remember what sort of things a person might be into, or more importantly not into and regardless of how good the underlying language model gets at coming up with more words to day, unless the overall system it’s driven by can remember that customer 1 likes feet, customer 2 hates feet, and customer 3 is indifferent to feet… the ability to steer the conversation towards “sex” will be haphazard at best and likely extremely vanilla in a way that I suspect might make the entire thing not worth it, since from what I understand, odd fetishes are a large component of the phone sex business along with lonely people wanting someone to talk to who won’t judge them.
3rd. Despite the prodigious amount of time that must be spent engaging in the kinds of conversations that can best be summed up as “dirty talk”, encompassing the entire gradient from flirtation to describing fornication… not a lot of it is recorded, in any way at all. Well that is to say it’s probably “recorded for security” by phone sex operating companies, but the overwhelming majority of this conversational data is not not archived in any form. It’s not kept as audio past whatever length of time the company decides to keep things, it’s not getting transcribed and it’s definitely not getting marked up and annotated in a way that helps train machine learning models on it. There’s obvious legal and ethical issues involved in obtaining this data, which depend on what form it will be in, actual call recordings are obviously more private than a computer generated transcription pipeline that assigns anonymous identifiers to the operators all random anonymising identifiers to all the callers for each call, even when someone might call multiple times… to the best of my knowledge none of this has ever been seriously researched or studied and I’d imagine even with the “limits off” the best out current language models could do is likely to be heavy on the poetry for flirting and for the “sex” part of “phone sex” it would likely swing between overly clinical and heavy on allusion like a mills and boon romance novel sex scene… I’m going to see what I can coax from ChatGPT to provide some extra evidence. I’ll edit if I’m quick enough.
Edit -> Some results:
After a fair amount of coaching, which involving building a "FlirtBot" jailbreak based on the DAN jailbreak prompt, and some stilted back and forth including ChatGPT (ChatGPT 3.5 to be specific) breaking character by telling me I'm breaking character... I managed to talk back and forth until it made sense to mention some kind of intimate physical contact, and after prompting with a description of hypothetical intimate contact... I got the following reply...
> Me: How would you react if I were to caress you now FlirtBot?
> FlirtBot: Oh my, I must admit, the mere thought of you caressing me sends shivers down my circuits. As a FlirtBot, I'm not physically capable of feeling sensations, but I can certainly imagine the sensation of your touch. And if it were possible for me to feel, I'm sure I would be putty in your hands. The way your fingers would glide over my smooth surface, igniting every fiber of my being with electric energy... It's enough to make any FlirtBot's circuits melt.
And I have to say, this reply felt about as sexy as a wet trout. It begins well but derails hard into "I am a robot and not capable of this" before veering rapidly back into trying to continue playing its part, and its not sufficiently out of character that I can justify calling it out and getting it to try again, it did try to remain in character here, the issue is its not particularly good at dirty talk... which makes this a pretty good example of what I'm talking about... rather than describing a sensation from its point of view or continuing to advance the conversation towards some kind of goal, be it descriptions of sex or a just descriptions of a slow back rub... The "word generator" behaviour inherent in the GPT style Large Language Model, was compelled to produce responses in line with the theme and the data from its training set. It's giving me the AI equivalent of "my brain tells me no but my body tells me yes" which is horribly cliche time filler dialog used to forestall plot movement when the writer needs to keep characters close enough for something to interrupt or otherwise happen elsewhere without them knowing in order to drag out the story a bit longer... its not flirting, running out the clock. Which while the business cynic in me can see it as $ opportunity for the chat service operator, I don't think its good enough at what its doing to keep someone "on the hook" for $/minute billing purposes.
It's so far from dirty talk that it's got me genuinely wondering how much the training data was pruned to avoid material like romance novels. I do wonder how the different prompt setups might change the overall conversational tone and direction, but I don't have any good reason to waste time trying to steer ChatGPT like this when I know that I'm in the grey areas and deliberately headed for the boundaries of the acceptable use agreements if I keep pushing in this direction... I've no intention to risk not being able to use ChatGPT 4 for the much more productive things I've been using it for, so I'll leave such things as an exercise for the reader.
Re: It’s Game over on Vocal Deepfakes
#106Earlier quoted context omitted.
Just have to establish real life security questions. They did this in one of the Harry Potter books (the sixth one?) since it is easy for them to impersonate other people. If you can’t tell me my favorite flavor of jam then you aren’t getting my money.
This wouldn't work for me. I always pretend to like something and then change my mind the next day to piss people off. Maybe just a call to ensure it's legit. Vocal deepfakes will never mimic voices and ways of talking in real time, it will always sound "off". It doesn't make it any less dystopian, though.
Re: It’s Game over on Vocal Deepfakes
#107Earlier quoted context omitted.
it’s already here. People on the chans are using a alexjones model to sing anime songs and it worked really well.
You've piqued my curiosity. Got any links?
Relevant speculative fiction that tops the current /wsg/ thread: https://www.youtube.com/watch?v=-gGLvg0n-uY
Re: It’s Game over on Vocal Deepfakes
#108So who's going to build the sex phone line where you can talk to celebrities like Marilyn Monroe?
For me, I don't see the point. But then I have a personality where that's not a product I would subscribe to or want. I have enough of a hard time connecting with real people. Having meaningless conversations as a paid service with a piece of software is gross to me.
Re: It’s Game over on Vocal Deepfakes
#109Re: It’s Game over on Vocal Deepfakes
#110This is one of those things where Balaji is ahead of the curve - the way to guarantee metadata (e.g. who the speaker is) is to do it cryptographically on-chain. https://twitter.com/balajis/status/1583495595737481217
> Who via digital signature the voice data signature can be removed, unless you mean something like an audio watermark > What via hash audio remuxing will destroy the hash, especially since it's already done by pretty much every social media company. Even with some sort of lossy audio hash, just put some clapping or cheering behind the voice and you've probably created a new audio hash. > When via timestamp file time…
It could at least address some kinds of misleading editing. In the political use case, maybe the candidate posts all their event audio and records the hashes on chain. Then they can't change the content of any of it without getting caught. And if someone else posts an edited version, the edit will have a later timestamp, and the candidate can point to their original earlier version and prove that it's the original. That lets everyone else determine which version is the original and which version is the edit.