NaturalSpeech 2: Zero-shot speech and singing synthesizers
speechresearch.github.io
NaturalSpeech 2: Zero-shot speech and singing synthesizers
1–10 of 126 posts
Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers
#2SCENE:
T-800, speaking to John Connor in normal voice: "What's the dog's name?"
John Connor: "Max."
T-800, impersonating John, on the phone with T-1000: "Hey Janelle, what's wrong with Wolfie? I can hear him barking. Is he all right?"
T-1000, impersonating John's foster mother, Janelle: "Wolfie's fine, honey. Wolfie's just fine. Where are you?"
T-800 hangs up the phone and says to John in normal voice: "Your foster parents are dead."
--
Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers
#3Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers
#4Wow, that ethics statement at the end
Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers
#5Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers
#6That being said, I think it is only a matter of time before cyber criminals develop an end to end fully automated penetration system that registers domain names, writes emails, makes phone calls, finds money mules, runs social media accounts, etc. all with a single console to run it all. That is a scary prospect for humanity and new tools for authenticating human identity will be needed - fast.
Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers
#7Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers
#8Transformers and Diffusion Models seem to be leading the pack lately in many tasks. It’s cool how these models can be used in a variety of quite different contexts without changing much about the network architecture. That being said, I think it is only a matter of time before cyber criminals develop an end to end fully automated penetration system that registers domain names, writes emails, makes phone calls, finds…
Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers
#9This sparks my interest so much since the last few days I was wondering if it was possible to use diffusion models on spectrograms to do audio effects editing. Here is a paper submitted a couple of weeks ago doing just that. And the demo examples are exceptional.
I want all of this to start slowing down a bit so I have a chance to catch up. I was just watching Andrej Karpathy's excellent Zero to Hero syllabus [3] trying to wrap my head around LLMs and now I feel I absolutely must catch up on diffusion models.
1. https://arxiv.org/abs/2304.00830
Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers
#10Transformers and Diffusion Models seem to be leading the pack lately in many tasks. It’s cool how these models can be used in a variety of quite different contexts without changing much about the network architecture. That being said, I think it is only a matter of time before cyber criminals develop an end to end fully automated penetration system that registers domain names, writes emails, makes phone calls, finds…
Governments already maintain registers of legally operating businesses: there's no reason that registration should not also be issuing cryptographic certificates which verify all forms of outbound communication by that business including phone calls.
But despite telecom being almost end-to-end digital (i.e. digital to the box on the street pretty much), there's been no push to close the last 100m. "Phone lines" shouldn't exist anymore with packet switched networking: you should just dial a path against a business, which is verifies itself with TLS certificates linked to it's business registry.