It’s Game over on Vocal Deepfakes
91–100 of 122 posts
Re: It’s Game over on Vocal Deepfakes
#92I tried ElevenLabs and it is truly amazing. You don't need much audio to train it and the final result can be incredible. I shared some of the snippets with friends and relatives and after the intitial scare (same as John) we agreed that the outcome wouldn't be different than a few years ago... We had impersonators for decades, so what have stopped political parties to hire an impersonator to create fake audios? The…
My impression about these voice generators was they're only good for Youtube tutorials and such. It was 2 years ago tho, I wonder how things changed.
Re: It’s Game over on Vocal Deepfakes
#93Those fake AI voices are perfect tools for phishing campaigns. I wonder if ChatGPT combined with a voice generator can mount successful phishing campaigns. It looks like Cyberpunk future we never asked is already here.
Would it really be wise to run a criminal operation while piping all your requests through Microsoft servers, considering the history of the company collaborating with law enforcement?
Years and years of bad software practices, corporate cost cutting measures resulted in a planet scale cybersecurity mess.
Current hacking scene is more similar to 80s than early 2000s, it is extremely hard to catch advanced actors.
Re: It’s Game over on Vocal Deepfakes
#94Earlier quoted context omitted.
I'm sure it's going to get exponentially better, but right now ChatGPT sounds like one very vanilla guy. The other day I asked it to write me a conversation between David Foster Wallace and Ernest Hemingway, and it came out something like (paraphrasing): HEMINGWAY: I notice you write about political issues and philosophy. WALLACE: Thank you, I noticed you write about raw experience and manliness. HEMINGWAY: Well, wha…
Has ChatGPT actually read any books (apart from snippets on the Internet), or has that been avoided for copyright reasons? I wonder how long until someone pirates every ebook on Z-Library or SciHub and trains a language model off that. Imagine a language model that's read every book and every scientific paper.
There’s a lot of papers that discredit older papers or work or entire subfields of study and the current large language models would just all all of it and go “these are words” with no analytical reasoning applied.
These are predictive mechanisms not analytical ones and until we get a better handle on how we can make them “stop and think”… I’m just not sure how much benefit each larger dataset will add beyond “sounds more human” while it’s core problems of “still hallucinates” and “can’t do basic reasoning” remain.
Also one question I’ve got on this front is how much duplicate data is processed out of the input sets? Using LibGen as an example there’s a lot of books that get reprinted and uploaded multiple times in different formats… does having these remain in the data provide a “valuable bias” or is it something that needs removing? Do the people building large language models pre-process their data to de-duplicate out these kinds of things from their current data sources?
Re: It’s Game over on Vocal Deepfakes
#95Earlier quoted context omitted.
"Hey mom, it's X. Sorry to leave this voice mail, I got in a bit of a car accident and lost my phone. I need to get back home, could you send $$$$ to ####?"
Just have to establish real life security questions. They did this in one of the Harry Potter books (the sixth one?) since it is easy for them to impersonate other people. If you can’t tell me my favorite flavor of jam then you aren’t getting my money.
Re: It’s Game over on Vocal Deepfakes
#96Earlier quoted context omitted.
Conceivably, anyone who uses his or her voice for spoken word or singing could, in theory, cryptographically sign each master copy of the recording, and therefore certify that they were present and active in its creation. Of course any post-processing would lose this attestation. So it's practically useless, I suppose. But it would be a first-line of defense in copyright cases, at least, by being able to say "I did i…
This is a thing I've started at pretty hard. Embedding a signal that survives basic postprocessing isn't that hard, but doing it in such a way that it stays well below the noise floor and survives extensive editing is not easy. If anyone is interested I could be convinced to put my work on GitHub, highly doubtful that it will go commercial unless I get a huge raft of free time.
Re: It’s Game over on Vocal Deepfakes
#97Once the generation times get to real-time, I dunno what's going to happen. I follow the voice acting community a lot and this is a big existential threat level of worry in that community. But I also see some positives for the narrative voice field, but at the expense of actual actors. The latest sequel to a favorite audio book series has the professional narrator pronouncing different character names and town names…
On the upside, I can now narrate my own life and thoughts as Morgan Freeman.
Re: It’s Game over on Vocal Deepfakes
#98Earlier quoted context omitted.
"Hey mom, it's X. Sorry to leave this voice mail, I got in a bit of a car accident and lost my phone. I need to get back home, could you send $$$$ to ####?"
Just have to establish real life security questions. They did this in one of the Harry Potter books (the sixth one?) since it is easy for them to impersonate other people. If you can’t tell me my favorite flavor of jam then you aren’t getting my money.
Re: It’s Game over on Vocal Deepfakes
#99I've used a tool to make myself sound like a girl, it is freakily realistic and uncanny hearing what it produces. Chinese love scam rings are going to be all over this in a year, coupled with image/video generation
I might be a little biased since I have used a lot of text to speech tools over the years (I routinely listen to hours of TTS generated audio per week to read technical books and research papers that aren’t in audiobook format) and consequently developed a bit of an “ear” for all the ways various TTS engines sound robotic, mangle pronunciation, and generally fail to read how a human would.
I’d be very curious which tool or service you used that sounded “freakily realistic”, care to share?
I’ve been looking for a good one to test out if running the TTS audio through it sounds any more or less robotic, since their robotic prosody gets deeply tiring with particular kinds of text after a while.
Re: It’s Game over on Vocal Deepfakes
#100I tried ElevenLabs and it is truly amazing. You don't need much audio to train it and the final result can be incredible. I shared some of the snippets with friends and relatives and after the intitial scare (same as John) we agreed that the outcome wouldn't be different than a few years ago... We had impersonators for decades, so what have stopped political parties to hire an impersonator to create fake audios? The…
How does ElevenLabs perform for dialogues with emotions? Like for movies, TV shows and games? My impression about these voice generators was they're only good for Youtube tutorials and such. It was 2 years ago tho, I wonder how things changed.