Earlier quoted context omitted.
It will unfortunately undoubtedly be used for mass automation of scams but text AI (and pre-AI automation) have been used for that for many years as well. Doesn't really make sense to say "ok we should allow all forms of AI besides voice because of scams", I think. But yes, there needs to be some spreading of public awareness.
That's if you answer phone calls from numbers not already in your contacts. For me all such numbers go to voicemail and if the voice is of someone i know ill just call them directly. If you do any of the above you are looking to be scammed!
Crossing the uncanny valley of conversational voice
91–100 of 231 posts
Re: Crossing the uncanny valley of conversational voice
#92Miles gets Arrested: Sesame.ai https://youtu.be/cGMO2hRNnv0
Re: Crossing the uncanny valley of conversational voice
#93AI voice is an overwhelmingly harmful technology. It's biggest use will be to hurt people.
Re: Crossing the uncanny valley of conversational voice
#94This was almost worse though because it did feel like a rude person just interrupting instead of a dumb computer not being able to pick up normal social cues around when the person they're listening to has finished.
Re: Crossing the uncanny valley of conversational voice
#95It's good, but it still sounds fake to me, but in a different way. The voice itself sounds like a human, undoubtedly. But the cadence and the rhythm of speaking are off. It sounds like someone who isn't a podcaster trying to speak in the personality of a podcaster. It just sounds like someone trying too hard and speaking in an unnatural way.
Yes that is very specific, but that's what it sounds like to my ear.
Re: Crossing the uncanny valley of conversational voice
#96This was already posted here: https://news.ycombinator.com/item?id=43221377 but I’m really surprised at the lack of attention this model is getting. The responsiveness and apparent personality are pretty mind blowing. It’s similar to what OpenAI had initially demoed for advanced voice mode, at least for the voice conversation portion. The demo interactions are recorded, which is mentioned in their disclaimer under th…
I'm surprised by the lack of attention that Gemini 2.0 with native audio output got. They have a demo at https://youtu.be/qE673AY-WEI, which I think is really good too. The main problem with Google's model is that this audio output is not supported by the API, but you can try it at https://aistudio.google.com.
In general, text to speech is pretty good nowadays I think. For example, this is a little math video that I made a few days ago: https://www.youtube.com/watch?v=G1mvLrCfjFM with the (old) Google text to speech API. Honestly, I think the narration is better than I personally could have done. It's calm, well pronounced, and sounds relatively enthusiastic.
Re: Crossing the uncanny valley of conversational voice
#97This is so good that it's disarming. People are going to blabber everything to it, so we need a local private model. It's a lot to ask, I know. Incredible tech.
> Our models will be available under an Apache 2.0 license. ^ from the post https://github.com/SesameAILabs/csm is empty for now, but I imagine they'll be releasing it soon: https://x.com/_apkumar/status/1895492615220707723
Re: Crossing the uncanny valley of conversational voice
#98It's good, but it still sounds fake to me, but in a different way. The voice itself sounds like a human, undoubtedly. But the cadence and the rhythm of speaking are off. It sounds like someone who isn't a podcaster trying to speak in the personality of a podcaster. It just sounds like someone trying too hard and speaking in an unnatural way.
This is an interesting take, and I'd guess that the training data for this probably did use podcasts as a source. Getting very realistic / real world conversational training data for an ai would be hard. Only a subset of us appear on podcasts, radio or tv and probably all speak in a slightly artificial manner when we do.
Re: Crossing the uncanny valley of conversational voice
#99Re: Crossing the uncanny valley of conversational voice
#100It's good, but it still sounds fake to me, but in a different way. The voice itself sounds like a human, undoubtedly. But the cadence and the rhythm of speaking are off. It sounds like someone who isn't a podcaster trying to speak in the personality of a podcaster. It just sounds like someone trying too hard and speaking in an unnatural way.
This is an interesting take, and I'd guess that the training data for this probably did use podcasts as a source. Getting very realistic / real world conversational training data for an ai would be hard. Only a subset of us appear on podcasts, radio or tv and probably all speak in a slightly artificial manner when we do.