Live data from Hacker News

Pushing the frontiers of audio generation

deepmind.google

41–50 of 114 posts

Re: Pushing the frontiers of audio generation

#42
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

It's due to the histrionic mental epidemic that we are going through. A lot of people are just like that IRL. They cannot just say "the food was fine", it's usually some crap like "What on earth! These are the best cheese sticks I've had IN MY EN TI R E LIFE!".

“I’m OBSESSED with the dipping sauce. So good.”

Re: Pushing the frontiers of audio generation

#43
But what's the end goal and audience here? I don't believe people will resonate with robots making "um" and "ohs" because people usually resonate with an artist, a producer, a writer, a singer etc. A human layer with which people can empathize is essential. This can work as long as people are deceived and don't know there is no human behind it. If however i find out that a video is AI -generated i instantly lose interest in it. There are e.g. a lot of AI-generated architecture videos on youtube at the moment, i have never wanted to listen to one, because i know the emotions will be fake.

Re: Pushing the frontiers of audio generation

#44
post #16

Earlier quoted context omitted.

Totally agree. Maybe it’s just the clips they chose, but it feels overfit on the weird conversational elements that make it impressive? Like the “oh yeahs” from the other person when someone is speaking. It is cool to see that natural flow in a conversation generated by a model, but there’s waaaay too much of it in these examples to sound natural. And I say all that completely slackjawed that this is possible.

> Like the “oh yeahs” from the other person when someone is speaking. I bet that if you select a British accent you will get fewer of them.

Cheeky bugger, you are

Re: Pushing the frontiers of audio generation

#45
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

We're used to hearing some kind of identity behind voices -- we unconsciously sense clusters of vocabulary, intonation patterns, ticks, frequent interruption vs quiet patience, silence tolerance, response patterns to various triggers, etc that communicate a coherent person of some kind. We may not know that a given speaker is a GenX Methodist from Wisconsin that grew up at skate parks in the suburbs, but we hear clus…

Whether this stops at the uncanny valley or progresses to specific "AI celebrity" voices, I'm left thinking the engineers involved in this never stopped to think carefully about whether this ought to be done in the first place.

Re: Pushing the frontiers of audio generation

#46
I think I put my finger on exactly why it sounds a bit uncanny-valley: it sounds like humans who are reading from a prepared 'bit' or 'script'.

We've all been on those webinars where it's clear -- despite the infusions (on cue) of "enthusiasm" from the speaker attempting to make it sound more natural and off-the-cuff -- that they are reading from a script.

It's a difficult-to-mask phenomenon for humans.

That all said, I actually have more grace for an AI sounding like this than I do for a human presenter reading from a script. Like, if I'm here "live" and paying attention to what you're saying, at least do me the service of truly being "here" with me and authentically communicating vs. simply reading something.

If you're going to simply read something, then just send it to me to read too - don't pretend it's a spontaneously synchronous communication.

Re: Pushing the frontiers of audio generation

#48
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

We're used to hearing some kind of identity behind voices -- we unconsciously sense clusters of vocabulary, intonation patterns, ticks, frequent interruption vs quiet patience, silence tolerance, response patterns to various triggers, etc that communicate a coherent person of some kind. We may not know that a given speaker is a GenX Methodist from Wisconsin that grew up at skate parks in the suburbs, but we hear clus…

To be fair, the majority of podcasts are from a group of generic white guys, and they almost sound identical to these AI generated ones. The AI actually seems to to do a better job too.

Re: Pushing the frontiers of audio generation

#49

Earlier quoted context omitted.

We're used to hearing some kind of identity behind voices -- we unconsciously sense clusters of vocabulary, intonation patterns, ticks, frequent interruption vs quiet patience, silence tolerance, response patterns to various triggers, etc that communicate a coherent person of some kind. We may not know that a given speaker is a GenX Methodist from Wisconsin that grew up at skate parks in the suburbs, but we hear clus…

To be fair, the majority of podcasts are from a group of generic white guys, and they almost sound identical to these AI generated ones. The AI actually seems to to do a better job too.

Citation absolutely needed. You call this fair?

> the majority of podcasts are from a group of generic white guys

Re: Pushing the frontiers of audio generation

#50

Earlier quoted context omitted.

We're used to hearing some kind of identity behind voices -- we unconsciously sense clusters of vocabulary, intonation patterns, ticks, frequent interruption vs quiet patience, silence tolerance, response patterns to various triggers, etc that communicate a coherent person of some kind. We may not know that a given speaker is a GenX Methodist from Wisconsin that grew up at skate parks in the suburbs, but we hear clus…

Whether this stops at the uncanny valley or progresses to specific "AI celebrity" voices, I'm left thinking the engineers involved in this never stopped to think carefully about whether this ought to be done in the first place.

"Surely my genAI product won't be used to spam zero-effort slop all over the internet!"

- guy whose genAI product will definitely be used to spam zero-effort slop all over the internet.

Post reply on HN