Live data from Hacker News

Pushing the frontiers of audio generation

deepmind.google

11–20 of 114 posts

Re: Pushing the frontiers of audio generation

#11
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

Totally agree. Maybe it’s just the clips they chose, but it feels overfit on the weird conversational elements that make it impressive? Like the “oh yeahs” from the other person when someone is speaking. It is cool to see that natural flow in a conversation generated by a model, but there’s waaaay too much of it in these examples to sound natural.

And I say all that completely slackjawed that this is possible.

Re: Pushing the frontiers of audio generation

#12
post #9

Is there a free (ad supported?) online tool without login that reads text that you paste into it? I often would like to listen to a blog post instead of reading it, but haven't found an easy, quick solution yet. I tried piping text through OpenAI's tts-1-hd, model and it is the first one I ever found that is human like enough for me to like listening to it. So I could write a tool for my own usecase that pipes the te…

There is on iOS. No ads. "Reader" by Eleven Labs. I haven't used it that much but have listened to some white papers and blogs (some of which were like 45 minutes) and it "just worked". Even let's you click text you want to jump to.

And it's Eleven Labs quality- which unless I've fallen behind the times is the highest quality TTS by a margin.

Re: Pushing the frontiers of audio generation

#13
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

That could just be the context though. Listening to a clip that's a demo of what the model can produce is very different to listening to a YouTube video that's using the model to generate speech about something you'd actually want to watch a video of.

Re: Pushing the frontiers of audio generation

#14
post #9

Is there a free (ad supported?) online tool without login that reads text that you paste into it? I often would like to listen to a blog post instead of reading it, but haven't found an easy, quick solution yet. I tried piping text through OpenAI's tts-1-hd, model and it is the first one I ever found that is human like enough for me to like listening to it. So I could write a tool for my own usecase that pipes the te…

I use ms edge for this exact use case. Works well enough on any platform

Re: Pushing the frontiers of audio generation

#15
post #10
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

Probably because you're expecting it and looking at a demo page. Put these voices behind a real video or advertisement and I would imagine most people wouldn't be able to tell that it's AI generated at all.

It'd be annoying to me whether it was AI or human. The faux-excitement and pseudo-bonhomie is grating. They should focus on how people actually talk, not on copying the vocal intonation of coked-up public radio presenters just back from a positive affirmation seminar.

Re: Pushing the frontiers of audio generation

#16
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

Totally agree. Maybe it’s just the clips they chose, but it feels overfit on the weird conversational elements that make it impressive? Like the “oh yeahs” from the other person when someone is speaking. It is cool to see that natural flow in a conversation generated by a model, but there’s waaaay too much of it in these examples to sound natural. And I say all that completely slackjawed that this is possible.

> Like the “oh yeahs” from the other person when someone is speaking.

I bet that if you select a British accent you will get fewer of them.

Re: Pushing the frontiers of audio generation

#17
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

> Example of a multi-speaker dialogue generated by NotebookLM Audio Overview, based on a few potato-related documents.

Listening to this on 1.75x speed is excellent. I think the generated speaking speed is slow for audio quality, bc it'd be much harder to slow-down the generated audio while retaining quality than vice versa.

Re: Pushing the frontiers of audio generation

#18
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

It's due to the histrionic mental epidemic that we are going through.

A lot of people are just like that IRL.

They cannot just say "the food was fine", it's usually some crap like "What on earth! These are the best cheese sticks I've had IN MY EN TI R E LIFE!".

Re: Pushing the frontiers of audio generation

#19
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

[deleted]

Re: Pushing the frontiers of audio generation

#20
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

Totally agree. Maybe it’s just the clips they chose, but it feels overfit on the weird conversational elements that make it impressive? Like the “oh yeahs” from the other person when someone is speaking. It is cool to see that natural flow in a conversation generated by a model, but there’s waaaay too much of it in these examples to sound natural. And I say all that completely slackjawed that this is possible.

I love the technology, but I really don't want AI to sound like this.

Imagine being stuck on a call with this.

> "Hey, so like, is there anything I can help you with today?"

> "Talk to a person."

> "Oh wow, right. (chuckle) You got it. Well, before I connect you, can you maybe tell me a little bit more about what problem you're having? For example, maybe it's something to do with..."

Post reply on HN