I tell friends that the scene below from T2 doesn't feel futuristic anymore. In fact, it now feels... almost mundane. I mean, a smart "script kiddie" with a bit of ML expertise can pull off this kind of deepfake voice spoofing on a relatively cheap desktop computer nowadays. We live in interesting times. SCENE: T-800, speaking to John Connor in normal voice: "What's the dog's name?" John Connor: "Max." T-800, imperso…
NaturalSpeech 2: Zero-shot speech and singing synthesizers
31–40 of 126 posts
Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers
#32Earlier quoted context omitted.
We're safe for a few months. Asked about how did T-800 figured the foster parents are dead based on the exchange, and got: "Yes, there is a strong cue in the exchange that suggests to the T-800 that John's foster parents are dead. The cue is that when the T-1000, impersonating Janelle, answers the phone and John asks about Wolfie, she responds by saying, "Wolfie's fine, honey. Wolfie's just fine." The use of the word…
You have to use gpt-4 for this stuff man. Direct Response from gpt-4: The T-800 figured out that John's foster parents were dead based on the exchange because when it asked about "Wolfie" (a made-up name for the dog), the T-1000, impersonating Janelle, did not correct the name and instead went along with it, saying "Wolfie's fine." If the real Janelle had been on the phone, she would have corrected the T-800 by stati…
Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers
#33This shouldn't be too surprising. Similar results were had with GPT-3 a while back, which is kind-of-able to produce audio or images, encoded as streams of tokens, when trained on that task, despite not being designed for it. A very interesting property was noted a few years ago by multiple researchers, I'm not sure who discovered it first. Transfer learning is unreasonably effective. If you were training an image ge…
https://arxiv.org/abs/2209.15162 https://llava-vl.github.io/
LLMs are already being grounded.
Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers
#34This shouldn't be too surprising. Similar results were had with GPT-3 a while back, which is kind-of-able to produce audio or images, encoded as streams of tokens, when trained on that task, despite not being designed for it. A very interesting property was noted a few years ago by multiple researchers, I'm not sure who discovered it first. Transfer learning is unreasonably effective. If you were training an image ge…
There's something a bit more mindblowing than that. Language models and vision models learn representations so similar that you can connect them with just a linear projection between image embedding and text embedding space(no training of the image encoder or llm required). https://arxiv.org/abs/2209.15162 https://llava-vl.github.io/ LLMs are already being grounded.
Relevant previous discussion:
Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers
#35Earlier quoted context omitted.
We're safe for a few months. Asked about how did T-800 figured the foster parents are dead based on the exchange, and got: "Yes, there is a strong cue in the exchange that suggests to the T-800 that John's foster parents are dead. The cue is that when the T-1000, impersonating Janelle, answers the phone and John asks about Wolfie, she responds by saying, "Wolfie's fine, honey. Wolfie's just fine." The use of the word…
You have to use gpt-4 for this stuff man. Direct Response from gpt-4: The T-800 figured out that John's foster parents were dead based on the exchange because when it asked about "Wolfie" (a made-up name for the dog), the T-1000, impersonating Janelle, did not correct the name and instead went along with it, saying "Wolfie's fine." If the real Janelle had been on the phone, she would have corrected the T-800 by stati…
Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers
#36Earlier quoted context omitted.
The problem here is authenticating that you are talking to a human and not an AI. Certificates don't help with that. Criminals will register temporary businesses and obtain certificates for them with no problem.
But those certs can come with reputations attached, and it prevents people from claiming that they're representatives of well known companies.
And you still don't know if your taking to a human or an AI.
Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers
#37Earlier quoted context omitted.
You have to use gpt-4 for this stuff man. Direct Response from gpt-4: The T-800 figured out that John's foster parents were dead based on the exchange because when it asked about "Wolfie" (a made-up name for the dog), the T-1000, impersonating Janelle, did not correct the name and instead went along with it, saying "Wolfie's fine." If the real Janelle had been on the phone, she would have corrected the T-800 by stati…
I used ChatGPT. Isn't it already based on GPT-4 since a few weeks ago? It's the "Mar 23" version.
alternatively, you can sign up/request for api access here - https://openai.com/waitlist/gpt-4-api
Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers
#38I tell friends that the scene below from T2 doesn't feel futuristic anymore. In fact, it now feels... almost mundane. I mean, a smart "script kiddie" with a bit of ML expertise can pull off this kind of deepfake voice spoofing on a relatively cheap desktop computer nowadays. We live in interesting times. SCENE: T-800, speaking to John Connor in normal voice: "What's the dog's name?" John Connor: "Max." T-800, imperso…
except that now the T-1000 will have access to the facebook or instagram of Janelle and will know all about Max
Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers
#39Earlier quoted context omitted.
We aren't even far off from an LLM being able to infer that the parents are fake on the basis of the dog name. I'm not even gonna touch the chain gun shooting up a parking lot aspect.
We're safe for a few months. Asked about how did T-800 figured the foster parents are dead based on the exchange, and got: "Yes, there is a strong cue in the exchange that suggests to the T-800 that John's foster parents are dead. The cue is that when the T-1000, impersonating Janelle, answers the phone and John asks about Wolfie, she responds by saying, "Wolfie's fine, honey. Wolfie's just fine." The use of the word…
Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers
#40Well. I'm sure that will take care of everything.