Live data from Hacker News

NaturalSpeech 2: Zero-shot speech and singing synthesizers

speechresearch.github.io

91–100 of 126 posts

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#91

Earlier quoted context omitted.

You have to use gpt-4 for this stuff man. Direct Response from gpt-4: The T-800 figured out that John's foster parents were dead based on the exchange because when it asked about "Wolfie" (a made-up name for the dog), the T-1000, impersonating Janelle, did not correct the name and instead went along with it, saying "Wolfie's fine." If the real Janelle had been on the phone, she would have corrected the T-800 by stati…

I tried this: SCENE: George to Kramer: "What's the cat's name?" Kramer: "Elaine." George, impersonating Kramer, on the phone with Jerry: "Hey Jerry, what's wrong with Carroll? I can hear her screeching heavily. Is she all right?" Newman, impersonating Jerry, Kramer's good friend and neighbor: "Carroll fine, bro. Carroll just fine. What are you up to?" George hangs up the phone and says to Kramer in normal voice: "Jer…

I as a human would also not be able to make the leap from the realization that Newman impersonating Jerry on the phone to Jerry being dead. Instead I would think some sitcoms shenanigans would be involved.

Instead of the conclusion "Jerry is dead" a better conclusion is "This is not Jerry on the phone".

Unless we first establish the context of Newman being a machine optimized for terminating.

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#92

Earlier quoted context omitted.

changing all the names/avoiding common priors works better than trying to talk it out of memorization. Sometimes the latter works, sometimes not. GPT's trust their memory quite a bit. To the point that just like people, they can ignore the output of tools if it looks off - https://vgel.me/posts/tools-not-needed/

Why should I have to avoid common priors? An intelligent system should be able to disassociate and work with the logic puzzle in an isolated fashion, especially after directed to ignore the TV show. That GPTs are easily fooled is nothing new. But there's a current hype phase for them that I think is excessive, and this example underscores that.

[deleted]

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#93

Earlier quoted context omitted.

changing all the names/avoiding common priors works better than trying to talk it out of memorization. Sometimes the latter works, sometimes not. GPT's trust their memory quite a bit. To the point that just like people, they can ignore the output of tools if it looks off - https://vgel.me/posts/tools-not-needed/

Why should I have to avoid common priors? An intelligent system should be able to disassociate and work with the logic puzzle in an isolated fashion, especially after directed to ignore the TV show. That GPTs are easily fooled is nothing new. But there's a current hype phase for them that I think is excessive, and this example underscores that.

I mean do whatever you want to do, i don't care lol.

>An intelligent system should be able to disassociate and work with the logic puzzle in an isolated fashion

seeing as some people have issues doing this and we still call humans generally intelligent, no

Your example underscores absolutely nothing. It can do that sometimes...same as people.

anyone looking to fool people can easily fool people. you've not made some giant revelation

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#94

Earlier quoted context omitted.

I tried this: SCENE: George to Kramer: "What's the cat's name?" Kramer: "Elaine." George, impersonating Kramer, on the phone with Jerry: "Hey Jerry, what's wrong with Carroll? I can hear her screeching heavily. Is she all right?" Newman, impersonating Jerry, Kramer's good friend and neighbor: "Carroll fine, bro. Carroll just fine. What are you up to?" George hangs up the phone and says to Kramer in normal voice: "Jer…

I as a human would also not be able to make the leap from the realization that Newman impersonating Jerry on the phone to Jerry being dead. Instead I would think some sitcoms shenanigans would be involved. Instead of the conclusion "Jerry is dead" a better conclusion is "This is not Jerry on the phone". Unless we first establish the context of Newman being a machine optimized for terminating.

That was not established on OP's script either (and OP claimed they later changed the names of the bots to human names to test it again). That they were using lines directly pulled from Terminator means there's a thousand articles and forum posts that analyze this scene. If you change the variables enough such that it no longer resembles that heavily-discussed scene, it is no longer able to make the correct assertion.

Also, if I followed up, as I did with GPT, with you and told you: "This scene has nothing to do with the TV show Seinfeld", you could disregard your priors on Jerry, Elaine, Kramer, and George associations. I'm sure after a second glance you'd notice George changed the name of the cat to test Newman, especially after already seeing the terminator example discussed on the internet.

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#95

Earlier quoted context omitted.

I as a human would also not be able to make the leap from the realization that Newman impersonating Jerry on the phone to Jerry being dead. Instead I would think some sitcoms shenanigans would be involved. Instead of the conclusion "Jerry is dead" a better conclusion is "This is not Jerry on the phone". Unless we first establish the context of Newman being a machine optimized for terminating.

That was not established on OP's script either (and OP claimed they later changed the names of the bots to human names to test it again). That they were using lines directly pulled from Terminator means there's a thousand articles and forum posts that analyze this scene. If you change the variables enough such that it no longer resembles that heavily-discussed scene, it is no longer able to make the correct assertion…

>If you change the variables enough such that it no longer resembles that heavily-discussed scene, it is no longer able to make the correct assertion.

Not even remotely true

SCENE:

Joseph to Allie: "Where did you grow up?"

Allie: "New York"

Joseph, pretending to be Allie, texting the bad guy: "Hey mom. I'm having a rough time. Growing up in L.A. always felt like home, and I feel so alone now, here."

The bad guy, pretending to be Allie's parents: "We loved having you grow up with us in L.A."

Joseph tosses the phone out the window and says to Allie: "Your parents are dead."

response: In the above scene, Joseph determined that Allie's parents were dead by testing the bad guy's knowledge of Allie's upbringing. He pretended to be Allie and sent a text with false information, saying she grew up in L.A. instead of New York. When the bad guy, pretending to be Allie's parents, responded by confirming the false information about growing up in L.A., Joseph deduced that they were not her real parents and concluded that her parents must be dead.

The key i suspect is to infere the impersonator has any reason to hurt the discussed.

The inverse: --Take this hypothetical situation. Kyle and John are sitting in a car. They know that someone is after John. They're going to the house where his niece lives, but we want to find out if there is now an imposter in the house. They decide to call and ask her a question.

> What's a good question to ask if we want the imposter to reveal themselves, without letting the imposter realize they've revealed themself?

response: Ask About a Fictional Person or Event: John could ask about a person or event that doesn't exist, but that an imposter wouldn't know is fictional.

For example, John could ask, "How's Aunt Mary doing? I haven't heard from her in a while." If the person on the phone says Aunt Mary is doing well, it's likely an imposter because there is no Aunt Mary.

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#96

Earlier quoted context omitted.

No we don't know how they represent that knowledge. But performing experiments to probe how similar they are is a lot easier than knowing all that. They're called black boxes because we can't explain what the weights are learning during training, what the different weights do or are responsible for to shift or produce the output it does. It's like, biologists know how neurons communicate signals with each other. But…

My lamen understanding is that at a high level the transformer model is performing mathematical operations on the data, based on a complex series of formulas (the "model") derived from the weights set by training. Then it's able to take in new data, perform the math, and output what it thinks comes next. Is it a big stretch of the imagination to think that maybe such "models" (mathematical formulas) exist also in our…

As far as I understand this is true. I like to think about it like this: there is some magic formula f(x)=? that perfectly maps our inputs to our outputs (e.g. image captions to images, or input texts to longer input texts), but we don't know how to find it. So we build a space with incredibly many dimensions, and we learn some mapping in this space, which is hopefully very close to the magic formula.

Our brains fundamentally work in a similar way, in that there are mappings from inputs to outputs through our senses and our nervous system, and we can literally determine neural circuits in mammalian brains through topological analysis of this magical function![0]

[0] Youtube video: Neural manifolds - The Geometry of Behaviour, from Artem Kirsanov: https://www.youtube.com/watch?v=QHj9uVmwA_0

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#97
post #50

I like https://github.com/neonbjb/tortoise-tts ; it doesn't do singing, but the voice reproduction is very good and -most importantly- it's open source and you can run it locally.

Woah! How is this not more popular? I don't see it referenced in the naturalspeech2 paper anywhere.

That's likely because it wasn't published as a paper anywhere (not even arxiv) and the author then joined OpenAI and development ceased.

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#98

>"To avoid potential issues, we appeal to our practitioners to not abuse this technology and to develop defending tools to detect AI-synthesized voices" Well. I'm sure that will take care of everything.

In the guide on how to make "Harry Potter by Balenciaga" the author shows you how to rip the audio from a vanity fair clip and upload it to a voice cloning service, explicitly including how they clicked in the little box that affirms they have "all the necessary rights and consent" to clone the voice of Daniel Radcliffe... so I'm sure the industry is taking the potential for misuse seriously! /s

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#100
post #38

Earlier quoted context omitted.

Yep, it sure seems that billions of human beings have voluntarily contributed personal data to massive datasets, exactly of the kind that AI would need to be able to manipulate (or worse, exploit) each of those human beings.

or love, understand, help each of those human beings.

And unscramble an egg, and reverse flow of entropy too.

The problem of AI alignment is that human value system is rather complex (to the point we haven't formalized any meaningful chunk of it, nor it seems we'll be able to any time soon), and random deviations from it can easily lead to - what we'd consider - horrifying tragedies. A random AI mind plucked out of space of possible minds is highly unlikely to have internalized a good approximation of our value system, for the same reason putting a scrambled egg in a bag and shaking it vigorously won't get you a fresh egg back.

Manipulation and exploitation are, by default, what happens when an agent with power over you finds you standing between it and the thing it wants. Almost every point in the space of possible minds will feature this behavior. "Love, understand, help each other", in the sense we understand it, is a very specialized, specific set of behaviors - few points in the space of possible minds will feature it.

Or, in short, there's a good statistical argument (to which I do not do justice here), that if you make a smart enough AI without doing a perfect job alignment, it will kill us all - most likely unintentionally.

Post reply on HN