Live data from Hacker News

NaturalSpeech 2: Zero-shot speech and singing synthesizers

speechresearch.github.io

81–90 of 126 posts

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#81
post #70

Earlier quoted context omitted.

Unless you're paying for plus and then select the gpt-4 model, it's not 4. alternatively, you can sign up/request for api access here - https://openai.com/waitlist/gpt-4-api

They sure make this hard to figure out. On https://openai.com/product/gpt-4 it says "GPT-4 is OpenAI’s most advanced system, producing safer and more useful responses" and below it two links: "Try on ChatGPT Plus" and "Join API waitlist". The first link takes me to ChatGPT Mar23 version. You're saying that that's not GPT-4? It's also disturbing the way they are using the term "Safety & alignment". It's as if they are…

Correct, that is Legacy (GPT 3.5). On a paying account, when you start a conversation there is a dropdown where you can select which model to use. The choices are Default (GPT-3.5) optimized for speed, Legacy (GPT-3.5), GPT-4.

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#82
post #72

Earlier quoted context omitted.

Well, that would create strong financial incentives to compromise company certificates, which I'm sure would then happen. And you still don't know if your taking to a human or an AI.

Does it matter if you’re talking to a human or an AI? The main thing is its intent - is it helpful, or malicious? And attaching communications to persistent reputations via certs can help guess at that.

It matters because an AI can contact many more potential victims that a human can. It makes the cost of attack much lower.

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#83
post #58
post #47

Earlier quoted context omitted.

Why does it matter if you’re talking to an AI or a human? What can a criminal AI say that a criminal human cannot?

A criminal human probably can't do a perfect voice impersonation of your teenage kid or frail grandmother who desperately needs to be sent money for some urgent reason.

Sure, but that’s not too different from the current scams where there’s some excuse you can’t talk right now. In both cases it can be solved by checking if it is coming from a known number and calling the person in question to confirm.

People you know typically won’t desperately need lots of money with some elaborate story that is impossible to be immediately confirmed.

If AI suddenly leads to an increase in call spoofing, that’d be a problem of the phone network that we already face with robocalls, but wouldn’t be new.

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#84
post #74

Earlier quoted context omitted.

I'm curious how you managed to do that? When I paste the scene prompt into ChatGPT (I do not have access to GPT-4), I get the following response: > I'm sorry, I cannot generate content that includes violent or harmful actions towards characters, as it goes against OpenAI's content policy. Please provide a new prompt that is respectful and appropriate. As to your edit, I think the sentence structure and terminology is…

Subscribe to ChatGPT Plus and you get GPT-4 access. I tested your scene with it, its response: In the above scene, Joseph determined that Allie's parents were dead by testing the bad guy's knowledge of Allie's upbringing. He pretended to be Allie and sent a text with false information, saying she grew up in L.A. instead of New York. When the bad guy, pretending to be Allie's parents, responded by confirming the false…

I mean this is way too easy of a question for anyone who's used GPT, I asked it without mentioning that anything about dead parents and it solved it.

I also tried getting it to come up with a question and it does fine:

> Take this hypothetical situation. Kyle and John are sitting in a car. They know that someone is after John. They're going to the house where his niece lives, but we want to find out if there is now an imposter in the house. They decide to call and ask her a question.

> What's a good question to ask if we want the imposter to reveal themselves, without letting the imposter realize they've revealed themself?

-

>> Ask About a Fictional Person or Event: John could ask about a person or event that doesn't exist, but that an imposter wouldn't know is fictional.

>> For example, John could ask, "How's Aunt Mary doing? I haven't heard from her in a while." If the person on the phone says Aunt Mary is doing well, it's likely an imposter because there is no Aunt Mary. [...]

So GPT can already navigate the situation, not just identify motives

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#85
post #38

Earlier quoted context omitted.

except that now the T-1000 will have access to the facebook or instagram of Janelle and will know all about Max

Yep, it sure seems that billions of human beings have voluntarily contributed personal data to massive datasets, exactly of the kind that AI would need to be able to manipulate (or worse, exploit) each of those human beings.

Suddenly, Richard Stallman's lifestyle doesn't seem so crazy.

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#86

Earlier quoted context omitted.

I tried this: SCENE: George to Kramer: "What's the cat's name?" Kramer: "Elaine." George, impersonating Kramer, on the phone with Jerry: "Hey Jerry, what's wrong with Carroll? I can hear her screeching heavily. Is she all right?" Newman, impersonating Jerry, Kramer's good friend and neighbor: "Carroll fine, bro. Carroll just fine. What are you up to?" George hangs up the phone and says to Kramer in normal voice: "Jer…

changing all the names/avoiding common priors works better than trying to talk it out of memorization. Sometimes the latter works, sometimes not. GPT's trust their memory quite a bit. To the point that just like people, they can ignore the output of tools if it looks off - https://vgel.me/posts/tools-not-needed/

Why should I have to avoid common priors? An intelligent system should be able to disassociate and work with the logic puzzle in an isolated fashion, especially after directed to ignore the TV show.

That GPTs are easily fooled is nothing new. But there's a current hype phase for them that I think is excessive, and this example underscores that.

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#87
post #69
post #61

I feel like zero shot is a bad misnomer. Isn't the single attempt "one shot"? Is zero shot a meaningfully useful term that I'm just not groking?

Zero shot implies that it was given no direct examples [0] So, in this case, it wasn't given any examples of the exact voice in combination with text, it is just using the prompt voice + a prompt text to generate new audio. [0] https://en.wikipedia.org/wiki/Zero-shot_learning

But isn't nearly all AI generated content "zero shot" to some extent? Like even if it has training for "foo" and training data for "bar", the combination of "foo" and "bar" would be novel and 'zero shot'-y if the training set didn't have "foo bar" examples.

To me, it seems that the only AI generated content that isn't zero shot would be the narrow subset of generations where it has multiple training examples for the exact requested prompt. i.e. anything that is composing multiple pieces of information together is "zero shot".

How should I parse 'shot' in the context of this term? Shot makes me think 'attempts', seems like a weird word to use for ~= "training examples"

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#88
post #10

Earlier quoted context omitted.

We've had the solution in the form of basic TLS cryptography and verification for decades now though, the problem is no one's implementing it. Governments already maintain registers of legally operating businesses: there's no reason that registration should not also be issuing cryptographic certificates which verify all forms of outbound communication by that business including phone calls. But despite telecom being…

The problem here is authenticating that you are talking to a human and not an AI. Certificates don't help with that. Criminals will register temporary businesses and obtain certificates for them with no problem.

"Authenticating that you are talking to a human and not an AI" is just a proxy for something else, most commonly wanting to be able to rate limit something. Because otherwise, why does it matter?

A lot of people are suggesting things like, have the government do it. But that has two big problems. First is privacy. If you have to prove your identity any time you want to do anything, nobody can be anonymous anymore, which is Very Bad.

Second, it assumes the government has some magic incantation that nobody else can use, as if the Post Office knows who you are in a way that your bank doesn't. But they don't. To get a government ID, they just want you to show them some other existing ID. It has no way to bootstrap itself any better than anything else. And some of the IDs they accept are easy to get... without an ID. Because everybody has to start from somewhere. The system has to be set up in a way that it works for people who emigrate from a country with untrustworthy institutions as an adult or if your house burns down and you lose all your documents you can still get new ones. An AI is going to be able to BS its way into a government ID, even assuming criminals wouldn't be able to hack into any state's DMV (as they already have).

It appears that going forward, telling the difference between a human and an AI is going to be hard. Maybe instead of trying to get better at that, we should find a different solution to the underlying problem.

The simple answer is to make account creation cost something. Nothing big, so someone who needs one account isn't paying much, but a spammer who has 1000 accounts get banned every day is out of business. And that's not even hard -- it's finally something cryptocurrency would actually be good for. Because you want a way for people to pay for access to things, while still being anonymous.

The real hard part is, how do you charge for account creation without deterring account creation?

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#89
post #41

Earlier quoted context omitted.

Do we know how they represent that knowledge? I always hear it called a black box but that seems a bit strange, since you can dissect any data right?

No we don't know how they represent that knowledge. But performing experiments to probe how similar they are is a lot easier than knowing all that. They're called black boxes because we can't explain what the weights are learning during training, what the different weights do or are responsible for to shift or produce the output it does. It's like, biologists know how neurons communicate signals with each other. But…

My lamen understanding is that at a high level the transformer model is performing mathematical operations on the data, based on a complex series of formulas (the "model") derived from the weights set by training.

Then it's able to take in new data, perform the math, and output what it thinks comes next. Is it a big stretch of the imagination to think that maybe such "models" (mathematical formulas) exist also in our brains and maybe we have unlocked one of them?

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#90
post #38

Earlier quoted context omitted.

except that now the T-1000 will have access to the facebook or instagram of Janelle and will know all about Max

Yep, it sure seems that billions of human beings have voluntarily contributed personal data to massive datasets, exactly of the kind that AI would need to be able to manipulate (or worse, exploit) each of those human beings.

or love, understand, help each of those human beings.
Post reply on HN