Live data from Hacker News

Voice Synthesis for in-the-Wild Speakers via a Phonological Loop

ytaigman.github.io

21–27 of 27 posts

Re: Voice Synthesis for in-the-Wild Speakers via a Phonological Loop

#21

To me this is very exciting. I'm already working on my own home digital assistant modeled as NeNe Leaks from the Real Housewives to add personality to otherwise boring conversations with a robot. I've been looking at various style transfer techniques, and having something a bit more plug & play will help me focus on the more unique parts. I predict that we'll see more celebrity voices used as conversational interface…

What language are you working in? I've been working in Powershell out of convenience, but am looking to port my speech bot to Node.

Re: Voice Synthesis for in-the-Wild Speakers via a Phonological Loop

#23
post #18

I still think emphasis on a word or syllable is important here as there is far more information than you realize being conveyed with inflection. Consider: I am going to eat the ham sandwich = Me, no one else I am going to eat the ham sandwich = Nothing can stop me I am going to eat the ham sandwich = On my way; got distracted I am going to eat the ham sandwich = In case you doubt my intent I am going to eat the ham s…

This made me chuckle, and it's a great illustration of how much meaning is changed by different emphasis. However, I would read the to example like this:

I am going to eat the ham sandwich = The sandwich is the reason I am going (to the party, or wherever...)

Re: Voice Synthesis for in-the-Wild Speakers via a Phonological Loop

#24
post #8

Similar in quality to Lyrebird https://soundcloud.com/user-535691776/dialog Google WaveNet sounds almost perfect in comparison: https://deepmind.com/blog/wavenet-generative-model-raw-audio...

Some of the generated speech clips are unsettlingly robotic while WaveNet sounds passable, but it's the piano compositions that I found unnerving. I can't explain how but randomly generated music sounds so hollow and cold.

Methinks we should compose a musical Turing test...

Re: Voice Synthesis for in-the-Wild Speakers via a Phonological Loop

#25

To me this is very exciting. I'm already working on my own home digital assistant modeled as NeNe Leaks from the Real Housewives to add personality to otherwise boring conversations with a robot. I've been looking at various style transfer techniques, and having something a bit more plug & play will help me focus on the more unique parts. I predict that we'll see more celebrity voices used as conversational interface…

What language are you working in? I've been working in Powershell out of convenience, but am looking to port my speech bot to Node.

I'm working in Go right now, but probably will end up with a mix of various things. What's nice about Go is that it's very portable, faster than Python, nicer RAM/storage usage than Node (not needing a JIT and all), and I can cross compile binaries and distribute them to Raspberry Pis or whatever.

Re: Voice Synthesis for in-the-Wild Speakers via a Phonological Loop

#26
post #14

Earlier quoted context omitted.

Subvocalisation. Your own voice speaking quietly in the background to something else. Ideal for consumer indoctrination.

The voice we hear in our head and the one everyone else hear is starkly different.

The voice I 'hear' in my head and the one I actually hear when talking is starkly different.

Re: Voice Synthesis for in-the-Wild Speakers via a Phonological Loop

#27
post #14

Earlier quoted context omitted.

Subvocalisation. Your own voice speaking quietly in the background to something else. Ideal for consumer indoctrination.

The voice we hear in our head and the one everyone else hear is starkly different.

So, ads spoken with your voice, as you would hear it from your head.
Post reply on HN