Live data from Hacker News

Crossing the uncanny valley of conversational voice

sesame.com

81–90 of 231 posts

Re: Crossing the uncanny valley of conversational voice

#81
It's very good, really impressive demo. My feedback would be, Maya needs to keep quiet a little longer after asking a question. She would ask something, then as I thought about my reply, already be on to the next thing. It left me with the impression she was a babbler (which is not an unrealistic model of how humans are, but it would be cool to be able to dial such traits up or down to taste).

I suppose the lack of visual cues probably hinders things in that regard.

Re: Crossing the uncanny valley of conversational voice

#82

It's good, but it still sounds fake to me, but in a different way. The voice itself sounds like a human, undoubtedly. But the cadence and the rhythm of speaking are off. It sounds like someone who isn't a podcaster trying to speak in the personality of a podcaster. It just sounds like someone trying too hard and speaking in an unnatural way.

To me the actual words it used also seemed fake, sort of too deliberately breezy.

Re: Crossing the uncanny valley of conversational voice

#83
post #8

I asked it if it could whisper, and it replied in full voice, ”I’m whispering to you right now”.

Yeah it's definitely going through text still. I tried to get it to sing a song so it output some lyrics and then read them as a poem.

I did manage to get it to output "la la la la" and then it kind of sang them with a random melody.

It also can't say things loud and its idea of whispering for me was to say "pst".

Still apart from that it's very impressive!

Re: Crossing the uncanny valley of conversational voice

#85

I would say most command and control voice interactions are going to be like buying a coffee — the parameters of the transaction are well known, so it’s just about fine tuning the match between what the user wants and what the robot has to do. A small minority of these interactions are going to be like a restaurant server — chit chat, pleasantries, some information gathering, followed by issuing direct orders. The tr…

The most immediate application for this might be in replacing call centers in various roles. And most of those are very conversational.

For example tech support is in large parts about making the caller feel heard and getting them to do trouble shooting steps without feeling stupid. Sales is in large parts about getting the right person to talk to you and to keep them talking to you.

Re: Crossing the uncanny valley of conversational voice

#86

It's very good, really impressive demo. My feedback would be, Maya needs to keep quiet a little longer after asking a question. She would ask something, then as I thought about my reply, already be on to the next thing. It left me with the impression she was a babbler (which is not an unrealistic model of how humans are, but it would be cool to be able to dial such traits up or down to taste). I suppose the lack of v…

I think part of the issue is for the latency to be as low as this they have to tune their speech to text to find endpoints in very small increments and then send the text to the model immediately.

So unless the system has a lot of engineering and/or training put into the main model being able to recognize exactly when it should keep waiting versus a real response, it will just see something like "user: empty response" or "user: uhmm" and assume it is supposed to respond to that.

Re: Crossing the uncanny valley of conversational voice

#87
Text-To-Speech models still aren't trained on rich enough data to have all the nuances we need to be fully expressive. For example, most models don't have a way to change accents separately from language (e.g. English with a slight French accent) or have an ability to set emotions such as excitement or sleepiness.

We aren't even talking about adding laughing, singing/rap or beatboxing.

Re: Crossing the uncanny valley of conversational voice

#88
post #22

The intelligence of the model is very low though. I asked it about catcalling and it started to talk about cats!

There is a limit due to the need to keep model responses nearly instant and the trade off that smaller models that are generally capable of that have. Unless you have unique hardware Only Cerebras can run medium to large models at truly near instant speed.
Post reply on HN