Live data from Hacker News

Crossing the uncanny valley of conversational voice

sesame.com

191–200 of 231 posts

Re: Crossing the uncanny valley of conversational voice

#191

Earlier quoted context omitted.

The same reason why text LLMs show exaggerated emotions (enthusiasm about your questions, super-apologetic tone when you dislike the answer, etc). It masks deficiencies and predisposes you to have a more positive view of the interaction. Think of the most realistic and immediate ways to monetize this tech. It's customer support. Replacing sprawling outsourced call centers with a chat bot that has access to a couple o…

Has it been tried the other way? I don't remember an iteration where they weren't obnoxiously over-endearing. After the initial novelty, it would be better to reduce the amount of fake information you have to read, and any attempt at pretending to be a human is completely fake information at this point.

You can always tell it to respond critically and it will. In fact, I've been doing this for quite a few queries after getting the bubbly endearing first pass, and it really strips the veil away (and often makes things more actionable)

Re: Crossing the uncanny valley of conversational voice

#192
post #43
post #5

AI voice is an overwhelmingly harmful technology. It's biggest use will be to hurt people.

Cue all the responses saying "it's already been possible to harm people, AI doesn't fundamentally change anything, nothing to worry about"

Counter point: We were barely doing anything about it when bad actors were pwning people pre-AI, like with social media propaganda or romance scams.

And if we still do nothing about it post-AI? Well, that is already the status quo, so caring now feels performative unless we're going to finally chit chat about solutions.

The same could be said for the internet. "The internet can be used for bad" is an empty, trivial claim, not an insight that needs a standing ovation. The conversation we need is what to do about it. And the solutions need to be real ones, not "we need to put the cat back in the bag".

Re: Crossing the uncanny valley of conversational voice

#193
post #118

Earlier quoted context omitted.

one thing: language learning

When I meet people in VR who are ESL, I can tell based on their accent and mannerisms that they learned English by playing video games with westerners or watched a lot of YouTube. Do we really want to dilute the uniqueness of language by making everyone sound like they came out of a lab in California?

>Do we really want to dilute the uniqueness of language

I can't speak to whether it's desirable or not, but this has been happening with the advent of radio, movies, and television for over a century. So, are we worse off now, linguistically-speaking, than then? Do we really even notice missing accents if we never grew up with them?

Re: Crossing the uncanny valley of conversational voice

#194
post #164

Tried to do the demo but it kept cutting every sentance off half way through. When I told it that I couldnt understand it because their voice kept cutting off, it said 'oh you noticed that did you? Sorry about that we are still working out some kinks' - all perfectly with no cutting out. I fail to see that as coincidence.

Try headphones.

Re: Crossing the uncanny valley of conversational voice

#195

Earlier quoted context omitted.

I would like to think the child is missing the bonding and fun the two of you enjoyed with the robot guy. The child may be missing the experience of being with you and the robot guy. I would look for more activities you can explore with the child.

Honestly, I think if I wasn't there, she still would have loved it. She related to it like a person.

That sounds dangerous to me. Not like I think you did something wrong or exposed your daughter to danger at all; it was probably a really useful exercise. The scary part to me is how readily she accepted it as human, or friendly.

We already know how well people are deceived by text and images. Imagine if they're getting phone or video calls from "people" who keep them company for hours at a time. Imagine if they're accustomed to it from an early age. The notion of dealing with a real, messy, rough on the edges, honest human being well become an intractable frustration.

Re: Crossing the uncanny valley of conversational voice

#196

Earlier quoted context omitted.

>They have a demo at https://youtu.be/qE673AY-WEI That's not a demo, that's a video. Anyone can make something like that in an afternoon with a couple friends and a microphone. Also, Google is known for putting out fake "demos", remember the Google Duplex scam?

Scam? Duplex worked.

[deleted]

Re: Crossing the uncanny valley of conversational voice

#197

Earlier quoted context omitted.

>They have a demo at https://youtu.be/qE673AY-WEI That's not a demo, that's a video. Anyone can make something like that in an afternoon with a couple friends and a microphone. Also, Google is known for putting out fake "demos", remember the Google Duplex scam?

Scam? Duplex worked.

I doesn't work today, let alone 6 years ago.

But good work defending your master.

Re: Crossing the uncanny valley of conversational voice

#198
post #166

Earlier quoted context omitted.

Are you not reasoning on lossy abstractions? I'm still not on board with the (seemingly prevalent) notion that LLM's can't reason. What's reasoning, anyway? I'm not actively advocating for any side, but the arguments against reasoning always felt very tautological to me.

The burden of proof is on the argument that they _are_ reasoning, and I have seen very little evidence that they do. It's also immediately clear to me when I look at the architecture of transformers that reasoning is not in the cards. I could be convinced otherwise if, again, someone showed me an indication of reasoning behavior. Since there is no such evidence and the systems theory approach tells me it does not rea…

> It's also immediately clear to me when I look at the architecture of transformers that reasoning is not in the cards.

I'm not saying that's incorrect, but thb that's exactly the tautology I was talking about!

Re: Crossing the uncanny valley of conversational voice

#199

Earlier quoted context omitted.

Yeah this is straight up creepy, and I also can't stand chatgpt saying "Lmao" and "Yeah". Keep it formal & robotic.

What ever did you tell ChatGPT so it responded with "lmao"? I told it that it should behave explicitly like a computer in the system prompt, sort of worked.

After multiple prompts and utterly garbage output: https://i.imgur.com/5aOARCV.png

I'm almost positive that some AI systems have a backend that analyzes the sentiment of your messages and if you threaten to cancel billing it will notice your defcon-1 sentiment and spin up some more powerful instances behind the scenes to tide you over.

This is actually much more stressful than working without any AI as I have to decompress from constantly verbally obliterating a robotic intern.

I'll try with the system prompt. Also love your username.

Re: Crossing the uncanny valley of conversational voice

#200
Are there any technical innovations here over Moshi, which invented some of the pieces they use for their model? The only comparison I see is they split the temporal and depthwise transformers on the zeroth RVQ codebook, whereas Moshi has a special zeroth level vector quantizer distilled from a larger audio model, with the intent to preserve semantic information.

EDIT: also Moshi started with a pretrained traditional text LLM

Post reply on HN