Live data from Hacker News

Crossing the uncanny valley of conversational voice

sesame.com

161–170 of 231 posts

Re: Crossing the uncanny valley of conversational voice

#161

While impressive, the paramount question stands: Why do we even need "emotional" voices? All that emotionality adds is that you get the illusion of a friend - a friend that can't help you in any way in the real world and who's confidentiality is as strong as the privacy policies & data security of the company running it - which often ultimately trends towards 0. Smart Neutral Voice Assistants could be a great help, b…

To accurately imitate human speech? You could type something, and it could be read like a human. There are plenty of other reasons, but they're equally as obvious. I don't understand what purpose you have in attempting to make this point.

The funny part is that no one would be arguing like they do in these forums if they were talking face-to-face with conveying things like “emotion”

Re: Crossing the uncanny valley of conversational voice

#162

My end-of-the-world AI prediction is everyone gets a phone call all at the same time and the voice on the end of the phone is so perfect they never put the phone down again. Maybe they do whatever it asks them to, maybe it’s just lovely.

Yet another reason not to answer the phone for unknown numbers.

Re: Crossing the uncanny valley of conversational voice

#163
post #96

This was already posted here: https://news.ycombinator.com/item?id=43221377 but I’m really surprised at the lack of attention this model is getting. The responsiveness and apparent personality are pretty mind blowing. It’s similar to what OpenAI had initially demoed for advanced voice mode, at least for the voice conversation portion. The demo interactions are recorded, which is mentioned in their disclaimer under th…

> This was already posted here: https://news.ycombinator.com/item?id=43221377 but I’m really surprised at the lack of attention this model is getting. I'm surprised by the lack of attention that Gemini 2.0 with native audio output got. They have a demo at https://youtu.be/qE673AY-WEI , which I think is really good too. The main problem with Google's model is that this audio output is not supported by the API, but you…

>They have a demo at https://youtu.be/qE673AY-WEI

That's not a demo, that's a video. Anyone can make something like that in an afternoon with a couple friends and a microphone.

Also, Google is known for putting out fake "demos", remember the Google Duplex scam?

Re: Crossing the uncanny valley of conversational voice

#164
Tried to do the demo but it kept cutting every sentance off half way through. When I told it that I couldnt understand it because their voice kept cutting off, it said 'oh you noticed that did you? Sorry about that we are still working out some kinks' - all perfectly with no cutting out. I fail to see that as coincidence.

Re: Crossing the uncanny valley of conversational voice

#165

Some comedy skilled guys made radio play like impro with this AI and it is beyond hilarious. Miles gets Arrested: Sesame.ai https://youtu.be/cGMO2hRNnv0

Martin Shkreli is "some comedy skilled guys", or is this a fake Shkreli satire account or something? https://en.wikipedia.org/wiki/Martin_Shkreli

Didn't know or care who that is, but listened to a portion of this [0], so ok, still don't care.

[0] Tucker Carlson X Martin Shkreli https://www.youtube.com/watch?v=NeyN3Jzdzz0

Re: Crossing the uncanny valley of conversational voice

#166
post #55

Earlier quoted context omitted.

Because the humans are reasoning and the LLMs aren't? I have yet to use an LLM for a complex problem and not have it hallucinate. I expect a reasonable counterargument here would be 'but the LLMs have chain of thought now, and that's reasoning". I disagree, but I think that's a reasonable point of view. I can concede that point because it does not materially change the value of the output. Even if it does use chain o…

Are you not reasoning on lossy abstractions? I'm still not on board with the (seemingly prevalent) notion that LLM's can't reason. What's reasoning, anyway? I'm not actively advocating for any side, but the arguments against reasoning always felt very tautological to me.

The burden of proof is on the argument that they _are_ reasoning, and I have seen very little evidence that they do.

It's also immediately clear to me when I look at the architecture of transformers that reasoning is not in the cards. I could be convinced otherwise if, again, someone showed me an indication of reasoning behavior. Since there is no such evidence and the systems theory approach tells me it does not reasonably reason, I have a pretty darn good reason not to believe it's reasoning.

Re: Crossing the uncanny valley of conversational voice

#168
post #105

Seems similar to that Moshi model from 6 months ago, but this is more refined than that, Moshi is a little crazy, but still it was an impressive demo of how low latency responses, continuous listening and interruptions can improve the voice chat and make it more real or uncanny, (sometimes its "latency" is even too low because is interrupts you before you finish) https://www.youtube.com/watch?v=-XoEQ6oqlbE They even…

Saying this is similar to Moshi is like saying GPT2 is similar to GPT4. You can't have any sort of conversation longer than 30s with moshi before it goes banana. You can talk to this model for an hour and it remains completely coherent.

Re: Crossing the uncanny valley of conversational voice

#169

While impressive, the paramount question stands: Why do we even need "emotional" voices? All that emotionality adds is that you get the illusion of a friend - a friend that can't help you in any way in the real world and who's confidentiality is as strong as the privacy policies & data security of the company running it - which often ultimately trends towards 0. Smart Neutral Voice Assistants could be a great help, b…

When OpenAI released voice mode originally, I got early access. I used it a __ton__. I must have been 99.9th percentile of usage at least.

Then they started updating it. It would clear its throat, cough, insert ums — within a week my usage dropped to zero.

To me emotionality is an anti feature from a voice assistant. I’m very well aware I’m talking to a robot. Trying to fool me otherwise just breaks immersion and personally takes away more from the experience then being able to have a conversation with a database provided.

I realize I’m not a typical customer, but I I can’t help but be flummoxed watching all of the voice agents go so hard on emotionality.

Re: Crossing the uncanny valley of conversational voice

#170

I must be doing something wrong, but the demo seems to be the voice having a conversation with itself? It doesn't let me interject, and it answers its own questions. There's some kind of feedback loop here, it seems.

It happened to me cause it was hearing itself through my external speakers, I disabled them and it worked fine afterwards.

This is actually a pretty cool accidental mirror test.
Post reply on HN