Live data from Hacker News

Pushing the frontiers of audio generation

deepmind.google

31–40 of 114 posts

Re: Pushing the frontiers of audio generation

#31
post #9

Is there a free (ad supported?) online tool without login that reads text that you paste into it? I often would like to listen to a blog post instead of reading it, but haven't found an easy, quick solution yet. I tried piping text through OpenAI's tts-1-hd, model and it is the first one I ever found that is human like enough for me to like listening to it. So I could write a tool for my own usecase that pipes the te…

There is on iOS. No ads. "Reader" by Eleven Labs. I haven't used it that much but have listened to some white papers and blogs (some of which were like 45 minutes) and it "just worked". Even let's you click text you want to jump to. And it's Eleven Labs quality- which unless I've fallen behind the times is the highest quality TTS by a margin.

There's also the built-in "Speak Selection" feature you can enable in the accessibility settings.

Re: Pushing the frontiers of audio generation

#32
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

It's like their training set was made up entirely of awkward podcaster banter.

At least 83% Leo Laporte.

Re: Pushing the frontiers of audio generation

#33
post #16

Earlier quoted context omitted.

Totally agree. Maybe it’s just the clips they chose, but it feels overfit on the weird conversational elements that make it impressive? Like the “oh yeahs” from the other person when someone is speaking. It is cool to see that natural flow in a conversation generated by a model, but there’s waaaay too much of it in these examples to sound natural. And I say all that completely slackjawed that this is possible.

> Like the “oh yeahs” from the other person when someone is speaking. I bet that if you select a British accent you will get fewer of them.

Gor blimey lad, that's the problem now innit???

Re: Pushing the frontiers of audio generation

#34
post #9

Is there a free (ad supported?) online tool without login that reads text that you paste into it? I often would like to listen to a blog post instead of reading it, but haven't found an easy, quick solution yet. I tried piping text through OpenAI's tts-1-hd, model and it is the first one I ever found that is human like enough for me to like listening to it. So I could write a tool for my own usecase that pipes the te…

There is on iOS. No ads. "Reader" by Eleven Labs. I haven't used it that much but have listened to some white papers and blogs (some of which were like 45 minutes) and it "just worked". Even let's you click text you want to jump to. And it's Eleven Labs quality- which unless I've fallen behind the times is the highest quality TTS by a margin.

Reader is on a pretty good path to a monthly subscription model. Great audio quality, large selection of voices, and support for long-form input text.

Re: Pushing the frontiers of audio generation

#35
post #25

YouTube videos are already infested with insufferable AI elevator background "music". Even some channels that were previously good are using it. On the bright side, you can stop watching these channels and have more time for serious things.

> AI elevator background "music".

What are some examples? I haven't encountered this.

Re: Pushing the frontiers of audio generation

#36
post #6
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

It sounds like every sentence is an ad read.

Yeah... It isn't that it doesn't sound like human speech... it just sounds like how humans speak when they are uncomfortable or reading prepared and they aren't good at it.

Re: Pushing the frontiers of audio generation

#37
post #20

Earlier quoted context omitted.

Totally agree. Maybe it’s just the clips they chose, but it feels overfit on the weird conversational elements that make it impressive? Like the “oh yeahs” from the other person when someone is speaking. It is cool to see that natural flow in a conversation generated by a model, but there’s waaaay too much of it in these examples to sound natural. And I say all that completely slackjawed that this is possible.

I love the technology, but I really don't want AI to sound like this. Imagine being stuck on a call with this. > "Hey, so like, is there anything I can help you with today?" > "Talk to a person." > "Oh wow, right. (chuckle) You got it. Well, before I connect you, can you maybe tell me a little bit more about what problem you're having? For example, maybe it's something to do with..."

That's how the DJ feature of Spotify talks and it's pretty jarring.

"How's it going. We're gonna start by taking you back to your 2022 favorites, starting with the sweet sounds of XYZ". There's very little you can tweak about it, the suggestions kinda suck, but you're getting a fake friend to introduce them to you. Yay, I guess..

Re: Pushing the frontiers of audio generation

#38
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

We're used to hearing some kind of identity behind voices -- we unconsciously sense clusters of vocabulary, intonation patterns, ticks, frequent interruption vs quiet patience, silence tolerance, response patterns to various triggers, etc that communicate a coherent person of some kind. We may not know that a given speaker is a GenX Methodist from Wisconsin that grew up at skate parks in the suburbs, but we hear clus…

Agreed. To me it sounds like bad voice-over actors reading from a script. So the natural parts of a conversation where you might say the wrong thing and step back to correct yourself are all gone. Impressive for sure.

Re: Pushing the frontiers of audio generation

#39
post #9

Is there a free (ad supported?) online tool without login that reads text that you paste into it? I often would like to listen to a blog post instead of reading it, but haven't found an easy, quick solution yet. I tried piping text through OpenAI's tts-1-hd, model and it is the first one I ever found that is human like enough for me to like listening to it. So I could write a tool for my own usecase that pipes the te…

Both windows and macos (the operating systems) have this built-in under accessibility. It’s worth a try and I use it sometimes when I want to read something while cooking.

Re: Pushing the frontiers of audio generation

#40
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

It's because it's probably trained with "professional audio", ads, movies, audiobooks, and not "normal people talking". Like the effect when diffusion was mostly trained with stock photos.
Post reply on HN