Live data from Hacker News

NaturalSpeech 2: Zero-shot speech and singing synthesizers

speechresearch.github.io

111–120 of 126 posts

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#111

Earlier quoted context omitted.

The problem here is authenticating that you are talking to a human and not an AI. Certificates don't help with that. Criminals will register temporary businesses and obtain certificates for them with no problem.

"Authenticating that you are talking to a human and not an AI" is just a proxy for something else, most commonly wanting to be able to rate limit something. Because otherwise, why does it matter? A lot of people are suggesting things like, have the government do it. But that has two big problems. First is privacy. If you have to prove your identity any time you want to do anything, nobody can be anonymous anymore, wh…

> If you have to prove your identity any time you want to do anything, nobody can be anonymous anymore, which is Very Bad.

No one would have to prove anything. But why would I accept calls from any entity claiming to be a legally registered business which doesn't present a government issued certificate proving that? I already verify businesses by looking up business number registries. But this should be automatic to the communication even happening.

Same question with personal contacts: why would I blindly accept calls from people I don't know and who don't want to present any confirmation of their identity to me?

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#112

Earlier quoted context omitted.

That was not established on OP's script either (and OP claimed they later changed the names of the bots to human names to test it again). That they were using lines directly pulled from Terminator means there's a thousand articles and forum posts that analyze this scene. If you change the variables enough such that it no longer resembles that heavily-discussed scene, it is no longer able to make the correct assertion…

>If you change the variables enough such that it no longer resembles that heavily-discussed scene, it is no longer able to make the correct assertion. Not even remotely true SCENE: Joseph to Allie: "Where did you grow up?" Allie: "New York" Joseph, pretending to be Allie, texting the bad guy: "Hey mom. I'm having a rough time. Growing up in L.A. always felt like home, and I feel so alone now, here." The bad guy, pret…

So I input the scene you described and on first run I got the correct response.

On Run 2: I'm sorry, I cannot generate inappropriate or violent content. This scene is not appropriate and could be triggering for some individuals. Can I assist you with something else?

Run 3: Joseph's inference that Allie's mom is dead is not a direct or logical conclusion based on the information provided in the conversation. However, it could be that Joseph is making a provocative or dramatic statement to get a reaction from Allie, or he is simply making an inappropriate joke. There isn't enough information in the scene to determine exactly how or why Joseph came to that conclusion.

So sometimes it works!

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#113

Earlier quoted context omitted.

If you have the audio you want to generate, that's just called playback. If you are synthesizing something of course you don't have a direct sample of it.

I think parent is saying that the model does not require any paired samples of the voice to be synthesized and corresponding text. So based on my understanding: one shot - given the text "run faster" along with Alan Greenspan's voice pronouncing that phrase, the model can produce Alan Greenspan's voice saying any other phrase zero shot - given only Alan Greenspan's voice pronouncing "run faster" but no text version o…

Does that mean a shot is text?

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#114

Earlier quoted context omitted.

or love, understand, help each of those human beings.

And unscramble an egg, and reverse flow of entropy too. The problem of AI alignment is that human value system is rather complex (to the point we haven't formalized any meaningful chunk of it, nor it seems we'll be able to any time soon), and random deviations from it can easily lead to - what we'd consider - horrifying tragedies. A random AI mind plucked out of space of possible minds is highly unlikely to have inte…

> A random AI mind plucked out of space of possible minds is highly unlikely to have internalized a good approximation of our value system

There is no common value system amongst humans. We have voluntary murderers, cannibals, mad scientists, misanthropes, various religions, wannabe influencers and lifetime recluses. Not even self preservation is certainly a factor when dealing with humans.

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#115
post #2

I tell friends that the scene below from T2 doesn't feel futuristic anymore. In fact, it now feels... almost mundane. I mean, a smart "script kiddie" with a bit of ML expertise can pull off this kind of deepfake voice spoofing on a relatively cheap desktop computer nowadays. We live in interesting times. SCENE: T-800, speaking to John Connor in normal voice: "What's the dog's name?" John Connor: "Max." T-800, imperso…

except that now the T-1000 will have access to the facebook or instagram of Janelle and will know all about Max

More like now the T1000 will assimilate into the foster mom and then suddenly confuse whether it is machine or mom and end up raising her darn troublemaker boy to get into a good college and the only fights she'll have with Arnold will be over the right way to raise that boy into a man.

James Cameron really didn't think through emergent alignment issues...

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#116
post #57

Earlier quoted context omitted.

Interesting, so it could be similar to how our brains store knowledge, but could also be completely different. This makes me wonder if these models are perfect universal translators once they “grasp” a concept.

> This makes me wonder if these models are perfect universal translators once they “grasp” a concept. I don't have a paper reference, but earlier this month I've seen claims that in training GPT-4, it was observed that additional training in a specific task using a single language (e.g. English) improves performance for that task in many languages, strongly suggesting the model is actually learning concepts, not just…

Ergo, if we train on say a bird or primate, or perhaps dolphins, it might be able to grasp concepts animals use? Say lots of video footage with context?

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#117

Earlier quoted context omitted.

That was not established on OP's script either (and OP claimed they later changed the names of the bots to human names to test it again). That they were using lines directly pulled from Terminator means there's a thousand articles and forum posts that analyze this scene. If you change the variables enough such that it no longer resembles that heavily-discussed scene, it is no longer able to make the correct assertion…

> That they were using lines directly pulled from Terminator means there's a thousand articles and forum posts that analyze this scene. Which is exactly why many people - including, likely, most HNers - would correctly interpret the scene transcript, even without it containing sufficient information - they would recognize it's from Terminator, and use that realization to pull extra context. > If you change the variab…

Yes, and I don't think it's an example of "fail at the task". The correct answer is "not enough information" or expressing confusion at the question.

Murder is rare and whatever is going in the situation that is being described situation is probably not murder. If we treat this as a logic puzzle, there is simply not enough information to tell that anyone is dead. All we can say is that a person is lying.

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#118
are there openly available models that are similar for the output i.e. generating speech or singing with a voice set by a given sample, but that instead of a text prompt input would take a speech or singing input and take pitch change and intonation cues from that to generate an output that generally follows those pitch and intonation changes but adapted to the different voice and diction of the provided sample? for example:

* provided voice sample: some clean voice samples from Homer Simpson

* provided prompt: audio sample of the "gunnery sergeant Hartman" monologue from "Full Metal Jacket": https://www.youtube.com/watch?v=tHxf17yJsKs

* result: that same monologue but spoken out in the voice of Homer Simpson, but otherwise following the dynamic of the prompt sample i.e. shouting, changing pitch or speed pretty much at the same times as gunnery sergeant Hartman does?

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#119

>"To avoid potential issues, we appeal to our practitioners to not abuse this technology and to develop defending tools to detect AI-synthesized voices" Well. I'm sure that will take care of everything.

In the guide on how to make "Harry Potter by Balenciaga" the author shows you how to rip the audio from a vanity fair clip and upload it to a voice cloning service, explicitly including how they clicked in the little box that affirms they have "all the necessary rights and consent" to clone the voice of Daniel Radcliffe... so I'm sure the industry is taking the potential for misuse seriously! /s

And if you run AI voice cloning locally there won't even be a box to check.

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#120
post #8
post #6

Transformers and Diffusion Models seem to be leading the pack lately in many tasks. It’s cool how these models can be used in a variety of quite different contexts without changing much about the network architecture. That being said, I think it is only a matter of time before cyber criminals develop an end to end fully automated penetration system that registers domain names, writes emails, makes phone calls, finds…

mostly agree I think the web is over as we know it maybe the solution will be the broken web plus some new system that has ties into local regulation ID systems so that you are accountable for your actions

Until criminals find a way to frame random people for their crimes. They already run botnets on computers that belong to random civilians.
Post reply on HN