Live data from Hacker News

NaturalSpeech 2: Zero-shot speech and singing synthesizers

speechresearch.github.io

101–110 of 126 posts

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#101
Anyone else think its an abuse of terminology to refer to speaker conditioning as "in-context learning" and "prompting" now?

Like, using a reference encoder to condition on an unseen target speaker has been around for 4 or 5 years now. These results are already cool without mystifying the results by calling this "prompting" or "icl"

Unless there is some subtle difference that warrants the new terminology?

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#102

Earlier quoted context omitted.

I as a human would also not be able to make the leap from the realization that Newman impersonating Jerry on the phone to Jerry being dead. Instead I would think some sitcoms shenanigans would be involved. Instead of the conclusion "Jerry is dead" a better conclusion is "This is not Jerry on the phone". Unless we first establish the context of Newman being a machine optimized for terminating.

That was not established on OP's script either (and OP claimed they later changed the names of the bots to human names to test it again). That they were using lines directly pulled from Terminator means there's a thousand articles and forum posts that analyze this scene. If you change the variables enough such that it no longer resembles that heavily-discussed scene, it is no longer able to make the correct assertion…

> That they were using lines directly pulled from Terminator means there's a thousand articles and forum posts that analyze this scene.

Which is exactly why many people - including, likely, most HNers - would correctly interpret the scene transcript, even without it containing sufficient information - they would recognize it's from Terminator, and use that realization to pull extra context.

> If you change the variables enough such that it no longer resembles that heavily-discussed scene, it is no longer able to make the correct assertion.

Same with humans. Switch it enough it doesn't resemble the Terminator scene, while still making it underspecified (not enough information about the nature and intent of the persons involved), and humans will fail at the task too.

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#103

Earlier quoted context omitted.

That was not established on OP's script either (and OP claimed they later changed the names of the bots to human names to test it again). That they were using lines directly pulled from Terminator means there's a thousand articles and forum posts that analyze this scene. If you change the variables enough such that it no longer resembles that heavily-discussed scene, it is no longer able to make the correct assertion…

>If you change the variables enough such that it no longer resembles that heavily-discussed scene, it is no longer able to make the correct assertion. Not even remotely true SCENE: Joseph to Allie: "Where did you grow up?" Allie: "New York" Joseph, pretending to be Allie, texting the bad guy: "Hey mom. I'm having a rough time. Growing up in L.A. always felt like home, and I feel so alone now, here." The bad guy, pret…

This is scarily impressive.

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#104

Earlier quoted context omitted.

No we don't know how they represent that knowledge. But performing experiments to probe how similar they are is a lot easier than knowing all that. They're called black boxes because we can't explain what the weights are learning during training, what the different weights do or are responsible for to shift or produce the output it does. It's like, biologists know how neurons communicate signals with each other. But…

My lamen understanding is that at a high level the transformer model is performing mathematical operations on the data, based on a complex series of formulas (the "model") derived from the weights set by training. Then it's able to take in new data, perform the math, and output what it thinks comes next. Is it a big stretch of the imagination to think that maybe such "models" (mathematical formulas) exist also in our…

My layman take on LLMs is that they map tokens to points in an absurdly high-dimensional vector space (on the order of tens to hundred of thousands dimensions). The training process shifts those points around to make the related tokens closer, which eventually ends up encoding pretty much any kind of relationship you could think of between the tokens, semantic or otherwise, as proximity in one or more dimensions. The latent space has enough dimensions to accommodate all those relationships, which is how even tasks which require complex understanding of abstract concepts still boil down to adjacency search in that space.

In other words: the LLM isn't learning algorithms, it's building a high-dimensional point cloud, where things related to each other are closer together.

Now, IIRC the visual model mentioned above works with sub-1000-dimensional latent space, which to me feels like not enough... space to fit generalized concepts in. But then, the prompts to txt2img and img2img models I saw seem more like additive modifiers, with individual tokens mostly independent of each other, so maybe that explanation still fits.

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#105
post #57

Earlier quoted context omitted.

No we don't know how they represent that knowledge. But performing experiments to probe how similar they are is a lot easier than knowing all that. They're called black boxes because we can't explain what the weights are learning during training, what the different weights do or are responsible for to shift or produce the output it does. It's like, biologists know how neurons communicate signals with each other. But…

Interesting, so it could be similar to how our brains store knowledge, but could also be completely different. This makes me wonder if these models are perfect universal translators once they “grasp” a concept.

> This makes me wonder if these models are perfect universal translators once they “grasp” a concept.

I don't have a paper reference, but earlier this month I've seen claims that in training GPT-4, it was observed that additional training in a specific task using a single language (e.g. English) improves performance for that task in many languages, strongly suggesting the model is actually learning concepts, not just words.

If that's the case, then I think we have indeed accidentally made a universal translator (limited to humans, though).

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#106
post #50

I like https://github.com/neonbjb/tortoise-tts ; it doesn't do singing, but the voice reproduction is very good and -most importantly- it's open source and you can run it locally.

Woah! How is this not more popular? I don't see it referenced in the naturalspeech2 paper anywhere.

unfortunately it can't be used for real-time use cases, but YourTTS is just as good and faster (except... non-commercial)

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#107
post #69
post #61

I feel like zero shot is a bad misnomer. Isn't the single attempt "one shot"? Is zero shot a meaningfully useful term that I'm just not groking?

Zero shot implies that it was given no direct examples [0] So, in this case, it wasn't given any examples of the exact voice in combination with text, it is just using the prompt voice + a prompt text to generate new audio. [0] https://en.wikipedia.org/wiki/Zero-shot_learning

If you have the audio you want to generate, that's just called playback. If you are synthesizing something of course you don't have a direct sample of it.

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#108
post #76
post #26

Earlier quoted context omitted.

But those certs can come with reputations attached, and it prevents people from claiming that they're representatives of well known companies.

Where does the reputation come from? Is it going to work like our already questionable anti-spam systems?

You don't need reputation systems: I can count on two hands the number of organizations that have any reason to be contacting me unsolicited.

Add a minor tie-in with my bank, and my reputation list would essentially be "the government, my bank, my doctor, my insurer + any company I've done monetary transactions with recently".

And realistically in a world with this system, this is all already implemented by my phone's contact list - no more phone numbers would mean that there's no reason to think that any organization would be contacting me from a totally unique (or anonymous) identity. Instead I'd just have a contact book whitelist entry for "bank.com.au" or whatever scheme we ended up with.

As it is right now, actual government services call people from Caller ID blocked numbers, and don't widely publish allowable contact origins. Which is ridiculous.

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#109
post #42

Earlier quoted context omitted.

Like the strong financial incentives to compromise certificates owned by banking websites? Securing voice communications the same as website communications is an awesome idea, and the fact that it’s possible to steal some piece of data and compromise it shouldn’t prevent us from moving in that direction.

It has happened, but there has to be a big pay off to make it worthwhile. Smaller companies would likely be more frequent targets given they would be easier to compromise. By the way, I'm not suggesting that the idea is useless. I'm just pointing out it isn't a panacea, and it still doesn't address the core problem raised in the article that you don't know if you are speaking to a human or not.

But the point is that the pay off itself is limited. Smaller companies have a much smaller pool of people who would have any reason to accept unsolicited calls.

This is evident in how scams work today: you either get explicitly targeted, or you dragnet using some service which everyone has i.e. a bank, the tax office, or a telecom company.

Compromising "Joe's BBQ Emporium" might be easier, but it's still (1) something you have to do (and maybe get caught then) and (2) the number of people who are going to pick up, or not immediately blacklist unsolicited calls from "Joe's BBQ Emporium" is tiny.

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#110
post #69

Earlier quoted context omitted.

Zero shot implies that it was given no direct examples [0] So, in this case, it wasn't given any examples of the exact voice in combination with text, it is just using the prompt voice + a prompt text to generate new audio. [0] https://en.wikipedia.org/wiki/Zero-shot_learning

If you have the audio you want to generate, that's just called playback. If you are synthesizing something of course you don't have a direct sample of it.

I think parent is saying that the model does not require any paired samples of the voice to be synthesized and corresponding text. So based on my understanding:

one shot - given the text "run faster" along with Alan Greenspan's voice pronouncing that phrase, the model can produce Alan Greenspan's voice saying any other phrase

zero shot - given only Alan Greenspan's voice pronouncing "run faster" but no text version of what was said, the model can produce Alan Greenspan's voice saying any other phrase

Post reply on HN