Live data from Hacker News

Pushing the frontiers of audio generation

deepmind.google

1–10 of 114 posts

Re: Pushing the frontiers of audio generation

#3

Try it out in the demo https://cloud.google.com/text-to-speech/?hl=en and in the API https://cloud.google.com/text-to-speech/docs/create-dialogue...

If I change the language in the demo, it removes all my text and replaces it with a template text. That's bad.

Re: Pushing the frontiers of audio generation

#4
While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

Re: Pushing the frontiers of audio generation

#5
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

Agreed. To be fair, I also get annoyed by fake/exaggerated expression from human podcasters.

Re: Pushing the frontiers of audio generation

#6
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

It sounds like every sentence is an ad read.

Re: Pushing the frontiers of audio generation

#7
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

I suppose it doesn't matter if it is a human, or a bot delivering the message, if the message is boring

Re: Pushing the frontiers of audio generation

#9
Is there a free (ad supported?) online tool without login that reads text that you paste into it?

I often would like to listen to a blog post instead of reading it, but haven't found an easy, quick solution yet.

I tried piping text through OpenAI's tts-1-hd, model and it is the first one I ever found that is human like enough for me to like listening to it. So I could write a tool for my own usecase that pipes the text to tts-1-hd and plays the audio. But maybe there is already something with a public web interface out there?

Re: Pushing the frontiers of audio generation

#10
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

Probably because you're expecting it and looking at a demo page. Put these voices behind a real video or advertisement and I would imagine most people wouldn't be able to tell that it's AI generated at all.
Post reply on HN