Live data from Hacker News

Voice Synthesis for in-the-Wild Speakers via a Phonological Loop

ytaigman.github.io

11–20 of 27 posts

Re: Voice Synthesis for in-the-Wild Speakers via a Phonological Loop

#11

There are too many example to do fraud with this to list here. One example: Not too long ago I still did the rather more important banking stuff with a quick phone call (couldn't be done entirely online).

Would certainly make the phishing scheme where "fake CEO sends real CFO an email request to send a bank wire" more successful. Send the email, follow up with a voice mail.

Re: Voice Synthesis for in-the-Wild Speakers via a Phonological Loop

#14
post #13
post #7

Earlier quoted context omitted.

Isn't this scary ?

With how much people hate the sound of their own voices, I think that would backfire on the advertiser.

Subvocalisation. Your own voice speaking quietly in the background to something else. Ideal for consumer indoctrination.

Re: Voice Synthesis for in-the-Wild Speakers via a Phonological Loop

#15
post #14
post #13

Earlier quoted context omitted.

With how much people hate the sound of their own voices, I think that would backfire on the advertiser.

Subvocalisation. Your own voice speaking quietly in the background to something else. Ideal for consumer indoctrination.

The voice we hear in our head and the one everyone else hear is starkly different.

Re: Voice Synthesis for in-the-Wild Speakers via a Phonological Loop

#17
post #13
post #7

Earlier quoted context omitted.

Isn't this scary ?

With how much people hate the sound of their own voices, I think that would backfire on the advertiser.

It they could make it sound like my voice sounds from inside my head, and not like the weirdo that other people apparently hear, it would be scary on several levels.

Re: Voice Synthesis for in-the-Wild Speakers via a Phonological Loop

#18
I still think emphasis on a word or syllable is important here as there is far more information than you realize being conveyed with inflection.

Consider:

I am going to eat the ham sandwich = Me, no one else

I am going to eat the ham sandwich = Nothing can stop me

I am going to eat the ham sandwich = On my way; got distracted

I am going to eat the ham sandwich = In case you doubt my intent

I am going to eat the ham sandwich = I will not be juggling it

I am going to eat the ham sandwich = The ultimate ham sandwich will be mine

I am going to eat the ham sandwich = Not turkey, not roast beef

I am going to eat the ham sandwich = Between two slices of bread is what I do

Re: Voice Synthesis for in-the-Wild Speakers via a Phonological Loop

#19
To me this is very exciting. I'm already working on my own home digital assistant modeled as NeNe Leaks from the Real Housewives to add personality to otherwise boring conversations with a robot. I've been looking at various style transfer techniques, and having something a bit more plug & play will help me focus on the more unique parts. I predict that we'll see more celebrity voices used as conversational interfaces become more common.

Part of the complexity is going from 'context-free phonemes' to actually modeling personality. Having some way for the voice to know how to embed emotion, and ideally contextually from the sentences themselves. NeNe is an interesting example as she adds so many non-verbal sounds to her dialog (bleeps and bloops and eye rolls that she translates into affected speech). That's part of what makes her NeNe, and a big part of the entertaining value. Pursuing that is what will bring style transfer to the next level... total personality emulation. I fantasize about basic animatronics that can move her head side to side, twirl, and literally give eye rolls.

If anyone wants to work on this with me, give me a ping @azinman on twitter. I've currently been thinking about this as an open source project, but still holding out options as I continue development. I've got a ton more ideas she's integrating into with my bleeding edge smart home, far more than just personality emulation (including what I believe to be a breakthrough in passive context-sensing.. the real key to making the smart home actually smart).

Re: Voice Synthesis for in-the-Wild Speakers via a Phonological Loop

#20
post #18

I still think emphasis on a word or syllable is important here as there is far more information than you realize being conveyed with inflection. Consider: I am going to eat the ham sandwich = Me, no one else I am going to eat the ham sandwich = Nothing can stop me I am going to eat the ham sandwich = On my way; got distracted I am going to eat the ham sandwich = In case you doubt my intent I am going to eat the ham s…

This has always been a vexing problem in TTS. Going from context-free to contextual. First the machine will need to actually be able to model it's own intentions, and then figure out how to verbally do that. Difficult problems without solving hard AI. Easier if you can pre-determine the world/responses, as most "conversational interfaces" currently do.
Post reply on HN