Live data from Hacker News

Crossing the uncanny valley of conversational voice

sesame.com

91–100 of 231 posts

Re: Crossing the uncanny valley of conversational voice

#91

Earlier quoted context omitted.

It will unfortunately undoubtedly be used for mass automation of scams but text AI (and pre-AI automation) have been used for that for many years as well. Doesn't really make sense to say "ok we should allow all forms of AI besides voice because of scams", I think. But yes, there needs to be some spreading of public awareness.

That's if you answer phone calls from numbers not already in your contacts. For me all such numbers go to voicemail and if the voice is of someone i know ill just call them directly. If you do any of the above you are looking to be scammed!

Oh yes ? The scambot will leave a distress message and a number in your voicemail, using the voice of a relative. You would know better but I guarantee old people will call the number and strike a convo with the virtual relative.

Re: Crossing the uncanny valley of conversational voice

#93
post #5

AI voice is an overwhelmingly harmful technology. It's biggest use will be to hurt people.

I unfortunately agree with you. Old people with confusion/dementia, schizoid types, or very naive persons will fall for shattering scams. And the consequences on their grasp on reality will be terrible.

Re: Crossing the uncanny valley of conversational voice

#94
Still suffers the same problem that all Voice Recognition seems to suffer; cannot reliably detect that the speaker has finished speaking.

This was almost worse though because it did feel like a rude person just interrupting instead of a dumb computer not being able to pick up normal social cues around when the person they're listening to has finished.

Re: Crossing the uncanny valley of conversational voice

#95

It's good, but it still sounds fake to me, but in a different way. The voice itself sounds like a human, undoubtedly. But the cadence and the rhythm of speaking are off. It sounds like someone who isn't a podcaster trying to speak in the personality of a podcaster. It just sounds like someone trying too hard and speaking in an unnatural way.

It sounds like someone who is doing a microphone test for something they just bought and hearing themself on a delay from the monitoring.

Yes that is very specific, but that's what it sounds like to my ear.

Re: Crossing the uncanny valley of conversational voice

#96

This was already posted here: https://news.ycombinator.com/item?id=43221377 but I’m really surprised at the lack of attention this model is getting. The responsiveness and apparent personality are pretty mind blowing. It’s similar to what OpenAI had initially demoed for advanced voice mode, at least for the voice conversation portion. The demo interactions are recorded, which is mentioned in their disclaimer under th…

> This was already posted here: https://news.ycombinator.com/item?id=43221377 but I’m really surprised at the lack of attention this model is getting.

I'm surprised by the lack of attention that Gemini 2.0 with native audio output got. They have a demo at https://youtu.be/qE673AY-WEI, which I think is really good too. The main problem with Google's model is that this audio output is not supported by the API, but you can try it at https://aistudio.google.com.

In general, text to speech is pretty good nowadays I think. For example, this is a little math video that I made a few days ago: https://www.youtube.com/watch?v=G1mvLrCfjFM with the (old) Google text to speech API. Honestly, I think the narration is better than I personally could have done. It's calm, well pronounced, and sounds relatively enthusiastic.

Re: Crossing the uncanny valley of conversational voice

#97
post #13

This is so good that it's disarming. People are going to blabber everything to it, so we need a local private model. It's a lot to ask, I know. Incredible tech.

> Our models will be available under an Apache 2.0 license. ^ from the post https://github.com/SesameAILabs/csm is empty for now, but I imagine they'll be releasing it soon: https://x.com/_apkumar/status/1895492615220707723

let's hope they stay true to that, but with a-16-z being their VC, I can't imagine there isn't an ultimately exploitive end game in it.

Re: Crossing the uncanny valley of conversational voice

#98

It's good, but it still sounds fake to me, but in a different way. The voice itself sounds like a human, undoubtedly. But the cadence and the rhythm of speaking are off. It sounds like someone who isn't a podcaster trying to speak in the personality of a podcaster. It just sounds like someone trying too hard and speaking in an unnatural way.

This is an interesting take, and I'd guess that the training data for this probably did use podcasts as a source. Getting very realistic / real world conversational training data for an ai would be hard. Only a subset of us appear on podcasts, radio or tv and probably all speak in a slightly artificial manner when we do.

I agree, I thinks it's probably very easy to find billions of hours of conversation on YouTube, but non of it is set to training data with a good transcript.

Re: Crossing the uncanny valley of conversational voice

#99
a lot of comments are dismissive of these generated convos because of out how obvious it is that these convos are generated. i feel like that's a high bar. you can tell that GTA5 is generated, but it's close enough to be fun. i imagine that's as close as we'll get with conversational AI

Re: Crossing the uncanny valley of conversational voice

#100

It's good, but it still sounds fake to me, but in a different way. The voice itself sounds like a human, undoubtedly. But the cadence and the rhythm of speaking are off. It sounds like someone who isn't a podcaster trying to speak in the personality of a podcaster. It just sounds like someone trying too hard and speaking in an unnatural way.

This is an interesting take, and I'd guess that the training data for this probably did use podcasts as a source. Getting very realistic / real world conversational training data for an ai would be hard. Only a subset of us appear on podcasts, radio or tv and probably all speak in a slightly artificial manner when we do.

When I commented on the unnatural cadence, it told me that it had been trained on podcasts, which does help explain the issue - some people tend to “live-edit” themselves when a conversation is being recorded, which leads to this staccato. It seems they need to find a better source of training date for more natural conversational speech.
Post reply on HN