Live data from Hacker News

NaturalSpeech 2: Zero-shot speech and singing synthesizers

speechresearch.github.io

31–40 of 126 posts

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#31
post #2

I tell friends that the scene below from T2 doesn't feel futuristic anymore. In fact, it now feels... almost mundane. I mean, a smart "script kiddie" with a bit of ML expertise can pull off this kind of deepfake voice spoofing on a relatively cheap desktop computer nowadays. We live in interesting times. SCENE: T-800, speaking to John Connor in normal voice: "What's the dog's name?" John Connor: "Max." T-800, imperso…

except that now the T-1000 will have access to the facebook or instagram of Janelle and will know all about Max

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#32
post #13

Earlier quoted context omitted.

We're safe for a few months. Asked about how did T-800 figured the foster parents are dead based on the exchange, and got: "Yes, there is a strong cue in the exchange that suggests to the T-800 that John's foster parents are dead. The cue is that when the T-1000, impersonating Janelle, answers the phone and John asks about Wolfie, she responds by saying, "Wolfie's fine, honey. Wolfie's just fine." The use of the word…

You have to use gpt-4 for this stuff man. Direct Response from gpt-4: The T-800 figured out that John's foster parents were dead based on the exchange because when it asked about "Wolfie" (a made-up name for the dog), the T-1000, impersonating Janelle, did not correct the name and instead went along with it, saying "Wolfie's fine." If the real Janelle had been on the phone, she would have corrected the T-800 by stati…

I see this so often. People try to show what GPT can or can’t do without explicitly stating the version. Since not many people have access to GPT-4 it’s usually safe to assume they didn’t use it. GPT-4 is a massive improvement over 3.5 when it come to any kind of non-trivial logic or inference task.

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#33
post #30

This shouldn't be too surprising. Similar results were had with GPT-3 a while back, which is kind-of-able to produce audio or images, encoded as streams of tokens, when trained on that task, despite not being designed for it. A very interesting property was noted a few years ago by multiple researchers, I'm not sure who discovered it first. Transfer learning is unreasonably effective. If you were training an image ge…

There's something a bit more mindblowing than that. Language models and vision models learn representations so similar that you can connect them with just a linear projection between image embedding and text embedding space(no training of the image encoder or llm required).

https://arxiv.org/abs/2209.15162 https://llava-vl.github.io/

LLMs are already being grounded.

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#34
post #30

This shouldn't be too surprising. Similar results were had with GPT-3 a while back, which is kind-of-able to produce audio or images, encoded as streams of tokens, when trained on that task, despite not being designed for it. A very interesting property was noted a few years ago by multiple researchers, I'm not sure who discovered it first. Transfer learning is unreasonably effective. If you were training an image ge…

There's something a bit more mindblowing than that. Language models and vision models learn representations so similar that you can connect them with just a linear projection between image embedding and text embedding space(no training of the image encoder or llm required). https://arxiv.org/abs/2209.15162 https://llava-vl.github.io/ LLMs are already being grounded.

I was about to say the same thing. Hi og_kalu!

Relevant previous discussion:

https://news.ycombinator.com/item?id=35598281

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#35
post #13

Earlier quoted context omitted.

We're safe for a few months. Asked about how did T-800 figured the foster parents are dead based on the exchange, and got: "Yes, there is a strong cue in the exchange that suggests to the T-800 that John's foster parents are dead. The cue is that when the T-1000, impersonating Janelle, answers the phone and John asks about Wolfie, she responds by saying, "Wolfie's fine, honey. Wolfie's just fine." The use of the word…

You have to use gpt-4 for this stuff man. Direct Response from gpt-4: The T-800 figured out that John's foster parents were dead based on the exchange because when it asked about "Wolfie" (a made-up name for the dog), the T-1000, impersonating Janelle, did not correct the name and instead went along with it, saying "Wolfie's fine." If the real Janelle had been on the phone, she would have corrected the T-800 by stati…

I used ChatGPT. Isn't it already based on GPT-4 since a few weeks ago? It's the "Mar 23" version.

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#36
post #26

Earlier quoted context omitted.

The problem here is authenticating that you are talking to a human and not an AI. Certificates don't help with that. Criminals will register temporary businesses and obtain certificates for them with no problem.

But those certs can come with reputations attached, and it prevents people from claiming that they're representatives of well known companies.

Well, that would create strong financial incentives to compromise company certificates, which I'm sure would then happen.

And you still don't know if your taking to a human or an AI.

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#37
post #35

Earlier quoted context omitted.

You have to use gpt-4 for this stuff man. Direct Response from gpt-4: The T-800 figured out that John's foster parents were dead based on the exchange because when it asked about "Wolfie" (a made-up name for the dog), the T-1000, impersonating Janelle, did not correct the name and instead went along with it, saying "Wolfie's fine." If the real Janelle had been on the phone, she would have corrected the T-800 by stati…

I used ChatGPT. Isn't it already based on GPT-4 since a few weeks ago? It's the "Mar 23" version.

Unless you're paying for plus and then select the gpt-4 model, it's not 4.

alternatively, you can sign up/request for api access here - https://openai.com/waitlist/gpt-4-api

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#38
post #2

I tell friends that the scene below from T2 doesn't feel futuristic anymore. In fact, it now feels... almost mundane. I mean, a smart "script kiddie" with a bit of ML expertise can pull off this kind of deepfake voice spoofing on a relatively cheap desktop computer nowadays. We live in interesting times. SCENE: T-800, speaking to John Connor in normal voice: "What's the dog's name?" John Connor: "Max." T-800, imperso…

except that now the T-1000 will have access to the facebook or instagram of Janelle and will know all about Max

Yep, it sure seems that billions of human beings have voluntarily contributed personal data to massive datasets, exactly of the kind that AI would need to be able to manipulate (or worse, exploit) each of those human beings.

Re: NaturalSpeech 2: Zero-shot speech and singing synthesizers

#39
post #13

Earlier quoted context omitted.

We aren't even far off from an LLM being able to infer that the parents are fake on the basis of the dog name. I'm not even gonna touch the chain gun shooting up a parking lot aspect.

We're safe for a few months. Asked about how did T-800 figured the foster parents are dead based on the exchange, and got: "Yes, there is a strong cue in the exchange that suggests to the T-800 that John's foster parents are dead. The cue is that when the T-1000, impersonating Janelle, answers the phone and John asks about Wolfie, she responds by saying, "Wolfie's fine, honey. Wolfie's just fine." The use of the word…

[deleted]
Post reply on HN