Live data from Hacker News

Crossing the uncanny valley of conversational voice

sesame.com

121–130 of 231 posts

Re: Crossing the uncanny valley of conversational voice

#121
post #118

While impressive, the paramount question stands: Why do we even need "emotional" voices? All that emotionality adds is that you get the illusion of a friend - a friend that can't help you in any way in the real world and who's confidentiality is as strong as the privacy policies & data security of the company running it - which often ultimately trends towards 0. Smart Neutral Voice Assistants could be a great help, b…

one thing: language learning

language learning also works fine without emotionality faking, and is depending more on authentic speech recognition (e.g. you want the model to notice if you mispronounce important words, not gloss over it and just continue babble as otherwise this will bite you in the ass in the real world) as well as the system's overall specific ability to generate a personal learning curriculum.

Re: Crossing the uncanny valley of conversational voice

#122
post #118

While impressive, the paramount question stands: Why do we even need "emotional" voices? All that emotionality adds is that you get the illusion of a friend - a friend that can't help you in any way in the real world and who's confidentiality is as strong as the privacy policies & data security of the company running it - which often ultimately trends towards 0. Smart Neutral Voice Assistants could be a great help, b…

one thing: language learning

When I meet people in VR who are ESL, I can tell based on their accent and mannerisms that they learned English by playing video games with westerners or watched a lot of YouTube.

Do we really want to dilute the uniqueness of language by making everyone sound like they came out of a lab in California?

Re: Crossing the uncanny valley of conversational voice

#123
Impressive, but I think this is missing two important things to not sound robotic – some atmosphere and space. During a real conversation, both partners are in some kind of a space, either in room, park, car or just on foot in the street. So the voice must have a little bit of reverb according to the space this voice is located in, and there must be some bits of background noise present from that same space. Even lip movement provides some tiniest background noises when you speak which contributes to making the sound real.

Re: Crossing the uncanny valley of conversational voice

#124
post #39

Earlier quoted context omitted.

Now here's a little thought experiment. What does the world look like in 5 years when everyone is talking to these things that are indistinguishable from a real person? They will be funnier, more compassionate, less judgemental, smarter, and superficially "better" in every respect.

I've been asking this for years. People keep saying stuff like "but you'll want the human touch." Really? So when was the last time you asked someone for directions? Personally, I'd rather google something or discuss with ChatGPT than make someone listen to me for an hour. And that someone has to be extremely knowledgeable about a lot of different topics! Even here. Would I rather converse with y'all and get downvote…

Regarding your last question - well who knows, for sure, but: chess between humans is alive and well after the computers became unbeatable by humans.

I recently listened to an interview with magnus Carlsen on Joe rogan and found the angle of computers helping humans to “better understand the game” (as he put it) and improving human play (for learning, not playing humans) to be very interesting.

Whether that extends to human conversation, who knows. I for one would love to have a “her”-like companion, not for romance but to have a highly intelligent and patient and knowledgeable conversation partner to develop ideas with and learn from, and endless other uses - I think it’d add a lot to my and other peoples lives. I guess I agree with you.

Re: Crossing the uncanny valley of conversational voice

#126
I played around. Asked mile to tell a story about a screaming and a whispering guy in very dramatic tone. It couldn't do it as expressively as the voice samples on the page. It was plain reading mostly. I could hear that this generation is text based. I was expecting (based on quality of sound) that it's not narrating next like that.

Example: it was saying "two dude-us" while trying to tell a melodramatic story. Which I assume was originally "two dude...s" or something.

Re: Crossing the uncanny valley of conversational voice

#128

I played with this last night with my four-year old daughter. We had fun with asking Miles to explain what bones are made of etc. Today, she asked "where has that robot guy gone?". Crying now because I won't let her talk to Miles anymore. She has already developed an emotional connection to it. Worrying indeed.

I would like to think the child is missing the bonding and fun the two of you enjoyed with the robot guy. The child may be missing the experience of being with you and the robot guy. I would look for more activities you can explore with the child.

Honestly, I think if I wasn't there, she still would have loved it. She related to it like a person.

Re: Crossing the uncanny valley of conversational voice

#129
post #118

Earlier quoted context omitted.

one thing: language learning

When I meet people in VR who are ESL, I can tell based on their accent and mannerisms that they learned English by playing video games with westerners or watched a lot of YouTube. Do we really want to dilute the uniqueness of language by making everyone sound like they came out of a lab in California?

Likewise will you be learning how to speak formally or informally.

Getting that wrong in some languages e.g. Korean can be offensive.

Re: Crossing the uncanny valley of conversational voice

#130
post #96

This was already posted here: https://news.ycombinator.com/item?id=43221377 but I’m really surprised at the lack of attention this model is getting. The responsiveness and apparent personality are pretty mind blowing. It’s similar to what OpenAI had initially demoed for advanced voice mode, at least for the voice conversation portion. The demo interactions are recorded, which is mentioned in their disclaimer under th…

> This was already posted here: https://news.ycombinator.com/item?id=43221377 but I’m really surprised at the lack of attention this model is getting. I'm surprised by the lack of attention that Gemini 2.0 with native audio output got. They have a demo at https://youtu.be/qE673AY-WEI , which I think is really good too. The main problem with Google's model is that this audio output is not supported by the API, but you…

How do I get to this in aistudio.google.com?
Post reply on HN