Live data from Hacker News

Crossing the uncanny valley of conversational voice

sesame.com

131–140 of 231 posts

Re: Crossing the uncanny valley of conversational voice

#131

Earlier quoted context omitted.

People here are hallucinating too. So many people making obviously wrong claims with full confidence, which you only notice when it’s about something you know a lot about yourself.

They do but we have for instance education to reduce their hallucinations in narrow fields of expertise. And a system of guardrails to only let educated people work in those fields to avoid harm.

The above person was comparing it to friends and random online comments, though. I wouldn't be surprised if AI is far more reliable than those.

Re: Crossing the uncanny valley of conversational voice

#132

I would say most command and control voice interactions are going to be like buying a coffee — the parameters of the transaction are well known, so it’s just about fine tuning the match between what the user wants and what the robot has to do. A small minority of these interactions are going to be like a restaurant server — chit chat, pleasantries, some information gathering, followed by issuing direct orders. The tr…

I think an important application will be enabling the robots to have these conversations with each other, in order to replace actors.

Re: Crossing the uncanny valley of conversational voice

#133

While impressive, the paramount question stands: Why do we even need "emotional" voices? All that emotionality adds is that you get the illusion of a friend - a friend that can't help you in any way in the real world and who's confidentiality is as strong as the privacy policies & data security of the company running it - which often ultimately trends towards 0. Smart Neutral Voice Assistants could be a great help, b…

Yes, there are many use cases where emotional voices are not needed, but that's not the point.

The core is not to have emotional voices, but to train neural networks to emulate emotions (not just for voices). Humans are very emotional beings, and if you want to communicate with them effectively, you will need the emotional layer. Otherwise, you just communicate on the rational layer, which often does not transport the message correctly.

Think of humans as 20% rational and 80% emotional.

And I say that as a person who believed for a long time that I was 80% rational and just 20% emotional ;-)

Re: Crossing the uncanny valley of conversational voice

#135
post #123

Impressive, but I think this is missing two important things to not sound robotic – some atmosphere and space. During a real conversation, both partners are in some kind of a space, either in room, park, car or just on foot in the street. So the voice must have a little bit of reverb according to the space this voice is located in, and there must be some bits of background noise present from that same space. Even lip…

Which is... annoying in voice interactions on the web. I purposefully set up my mic to avoid any echo and sound pretty direct like a radio host. Adding a simulated environment is less of a problem than getting a good baseline.

Re: Crossing the uncanny valley of conversational voice

#136
post #55
post #39

Earlier quoted context omitted.

I've been asking this for years. People keep saying stuff like "but you'll want the human touch." Really? So when was the last time you asked someone for directions? Personally, I'd rather google something or discuss with ChatGPT than make someone listen to me for an hour. And that someone has to be extremely knowledgeable about a lot of different topics! Even here. Would I rather converse with y'all and get downvote…

Because the humans are reasoning and the LLMs aren't? I have yet to use an LLM for a complex problem and not have it hallucinate. I expect a reasonable counterargument here would be 'but the LLMs have chain of thought now, and that's reasoning". I disagree, but I think that's a reasonable point of view. I can concede that point because it does not materially change the value of the output. Even if it does use chain o…

This is like saying:

Gogole is great for one thing: brainstorming, and brainstorming is only useful if you have no idea what to do in the first place. Once you know _anything_ substantial about the subject matter Google loses its value to you.

Re: Crossing the uncanny valley of conversational voice

#137
Well I'm astounded. I talked to it for 13min, it crashed, but remembered the context when I returned a few minutes later and talked for a full 30min (it's limit).

It 99.9% felt like it performed at the level of Samantha in the movie Her.

I started asking all kinds of questions about how it worked and it mentioned a word I had to have it repeat because I hadn't heard it before: PROSODY (linguistics) — the study of elements of speech, including intonation, stress, rhythm and loudness, that occur simultaneously with individual phonetic segments: vowels and consonants. I asked about personality settings, à la TARS from Interstellar, and it said it automatically tailored responses by listening for tone and content.

It felt like the most "the future's here but not evenly distributed" interaction I've had since multi-touch on an original iPhone.

Re: Crossing the uncanny valley of conversational voice

#138
post #96

Earlier quoted context omitted.

> This was already posted here: https://news.ycombinator.com/item?id=43221377 but I’m really surprised at the lack of attention this model is getting. I'm surprised by the lack of attention that Gemini 2.0 with native audio output got. They have a demo at https://youtu.be/qE673AY-WEI , which I think is really good too. The main problem with Google's model is that this audio output is not supported by the API, but you…

How do I get to this in aistudio.google.com?

I think the one under "Stream Realtime" should be similar to the demo. It's only Gemini 2.0 flash though and not the full one.

Re: Crossing the uncanny valley of conversational voice

#139

I played with this last night with my four-year old daughter. We had fun with asking Miles to explain what bones are made of etc. Today, she asked "where has that robot guy gone?". Crying now because I won't let her talk to Miles anymore. She has already developed an emotional connection to it. Worrying indeed.

I can see how it's worrying, but mostly as a replacement for real connections - if instead it supplements them, then not so bad.

Most children love talking to a fun adult who enjoys talking to them. As parents we hope to be that adult for them most of the time, but of course that's not easy to do all the time.

If parents made a tool like this a crutch and it replaced quality time with them or they were less likely to hang out with their friends, then yeah that's a big problem. If they use it as a learning aide or occasional fun diversion, it seems great.

Re: Crossing the uncanny valley of conversational voice

#140
post #123

Impressive, but I think this is missing two important things to not sound robotic – some atmosphere and space. During a real conversation, both partners are in some kind of a space, either in room, park, car or just on foot in the street. So the voice must have a little bit of reverb according to the space this voice is located in, and there must be some bits of background noise present from that same space. Even lip…

Which is... annoying in voice interactions on the web. I purposefully set up my mic to avoid any echo and sound pretty direct like a radio host. Adding a simulated environment is less of a problem than getting a good baseline.

I think every microphone will give you some characteristic atmosphere and space for the voice recorded, so it's kind of a part of a sound baseline. It's only annoying when there is too much, but when it's only on the edge of perceivable it adds that naturality to the sound. You can reduce it to the minimum of course, but you cannot completely eliminate it. That slight room tone or mic signature kind of glues everything together, making it feel more real.
Post reply on HN