Live data from Hacker News

Crossing the uncanny valley of conversational voice

sesame.com

171–180 of 231 posts

Re: Crossing the uncanny valley of conversational voice

#171
post #96

Earlier quoted context omitted.

> This was already posted here: https://news.ycombinator.com/item?id=43221377 but I’m really surprised at the lack of attention this model is getting. I'm surprised by the lack of attention that Gemini 2.0 with native audio output got. They have a demo at https://youtu.be/qE673AY-WEI , which I think is really good too. The main problem with Google's model is that this audio output is not supported by the API, but you…

>They have a demo at https://youtu.be/qE673AY-WEI That's not a demo, that's a video. Anyone can make something like that in an afternoon with a couple friends and a microphone. Also, Google is known for putting out fake "demos", remember the Google Duplex scam?

Scam? Duplex worked.

Re: Crossing the uncanny valley of conversational voice

#172

I tried the demo, but I decided to not say anything. It desperately tried to make me talk. The entire experience was bizarre and unsettling - another commenter described it as a northern Californian startup CEO’s level of strange fake enthusiasm. As a Brit, I found the level of synthetic bubbliness in the voice extremely off-putting. I’d hate to live in a world where that was the way everyone behaved in real life. Th…

Will definitely need to tone down the American Corporate Alacrity for the UK market..

Re: Crossing the uncanny valley of conversational voice

#173

Hey, it’s Brendan from Sesame. The feedback is spot on. We still have so much to do to make it good. Inspiring but still many steps away from a great experience. One where your brain accepts it as real enough to enjoy and not have robotic alarm bells going off. Today, we’re firmly in the valley, but we’re optimistic we can climb out. Verbal communication is complex. There’s a big list of interesting challenges to tac…

Congrats, you invented hollywood style AGI in the eyes of many.

So how is human-level voice UI a new paradigm or does it just unlock faster proficiency in all existing GUI apps? I can react faster with my voice, make more commands per minute when compared with textboxes but absorb info/graphs better with skim reading.

Re: Crossing the uncanny valley of conversational voice

#174
I have so many questions. Is the model running client side? I was expecting to see webrtc used to send audio to a backend service, but instead i think i the audio waveform processing is done client side? Is it sending audio tokens over websockets to a backend service that is hosting the model? 1/16 slices are enough to accurately be able to recreate an audible sentence? Or is a speech to text model also running client side and are both text and tokens being sent to backend service? Is the backend sending audio tokens back or just text , with the text to speech running 100% client side? Is this using mimi codec or facebook's encodec?

Re: Crossing the uncanny valley of conversational voice

#175

While impressive, the paramount question stands: Why do we even need "emotional" voices? All that emotionality adds is that you get the illusion of a friend - a friend that can't help you in any way in the real world and who's confidentiality is as strong as the privacy policies & data security of the company running it - which often ultimately trends towards 0. Smart Neutral Voice Assistants could be a great help, b…

To accurately imitate human speech? You could type something, and it could be read like a human. There are plenty of other reasons, but they're equally as obvious. I don't understand what purpose you have in attempting to make this point.

[deleted]

Re: Crossing the uncanny valley of conversational voice

#176

I tried the demo, but I decided to not say anything. It desperately tried to make me talk. The entire experience was bizarre and unsettling - another commenter described it as a northern Californian startup CEO’s level of strange fake enthusiasm. As a Brit, I found the level of synthetic bubbliness in the voice extremely off-putting. I’d hate to live in a world where that was the way everyone behaved in real life. Th…

Will definitely need to tone down the American Corporate Alacrity for the UK market..

Just get rid of it all together. I want my device to sound dry and factual like the ship computer in Star Trek, not emotional and... moist... like the lovechild of a Youtuber and a SV startup bro.

Re: Crossing the uncanny valley of conversational voice

#179
I tried both models. I could easily tell Maya was AI, but Miles sounded so lifelike that I felt that initial apprehension like hopping on a conference line with strangers. I even chuckled at one of his side remarks. It was strange knowing it wasn’t a real person, but it was very hard not to feel like it was.

Re: Crossing the uncanny valley of conversational voice

#180
This is incredibly impressive. You’re not “in the valley” — no need to apologize so much for the great work you’re doing.

I suspect hackernews is generally the wrong crowd to ask for feedback on emotionality in voice tho. Some of these folks would prefer humans speak like robots.

Post reply on HN