Live data from Hacker News

Pushing the frontiers of audio generation

deepmind.google

101–110 of 114 posts

Re: Pushing the frontiers of audio generation

#101
post #43

But what's the end goal and audience here? I don't believe people will resonate with robots making "um" and "ohs" because people usually resonate with an artist, a producer, a writer, a singer etc. A human layer with which people can empathize is essential. This can work as long as people are deceived and don't know there is no human behind it. If however i find out that a video is AI -generated i instantly lose inte…

[deleted]

Re: Pushing the frontiers of audio generation

#102
post #43

But what's the end goal and audience here? I don't believe people will resonate with robots making "um" and "ohs" because people usually resonate with an artist, a producer, a writer, a singer etc. A human layer with which people can empathize is essential. This can work as long as people are deceived and don't know there is no human behind it. If however i find out that a video is AI -generated i instantly lose inte…

> I don't believe people will resonate with robots making "um" and "ohs" because people usually resonate with an artist, a producer, a writer, a singer etc.

I think they absolutely will, because "resonating" is not a material phenomenon, it's something people decide that they're doing. Your connection with an actor on television is not an actual connection. Most of acting is learning the times and length to be silent while making a particular face (dictated by the director) in order for the audience to project feelings and thoughts onto you. You're thinking about your camera blocking, or your groceries, and your audience sees you thinking about some plot point in a fictional world.

I've got a theory that we severely damaged a generation of girls by inundating them with images of girls their own age singing songs and acting parts all written and directed by middle-aged men - ones who chose as a profession to write songs in the voices of, write fiction in the voices of, and to direct, photograph and choreograph in person, tween girls. Their models of themselves have come from looking at these depictions of girls, who were never allowed to speak for themselves, and resonating.

Re: Pushing the frontiers of audio generation

#103

Earlier quoted context omitted.

We're used to hearing some kind of identity behind voices -- we unconsciously sense clusters of vocabulary, intonation patterns, ticks, frequent interruption vs quiet patience, silence tolerance, response patterns to various triggers, etc that communicate a coherent person of some kind. We may not know that a given speaker is a GenX Methodist from Wisconsin that grew up at skate parks in the suburbs, but we hear clus…

To be fair, the majority of podcasts are from a group of generic white guys, and they almost sound identical to these AI generated ones. The AI actually seems to to do a better job too.

I laughed and sort of agreed with this, in spirit. You're off on the details I think though.

When I listened to the audio samples before coming to the comments, I thought: "oh, like those totally lifeless and bland U.S. accents from podcasts, YT, etc."

I wouldn't associate it with skin colour or gender though at all. I've no idea why you'd go there - any skin colour and any gender is absolutely welcomed into the fold of U.S. cultural production, if they can produce bland generic "content" sincerely enough, it seems to me.

Disclaimer: many U.S. accents are interesting and wonderful (Colorado; Tom Waits), they don't all sound generic and bland. I have U.S. friends therefore I can pass judgment (TM).

Re: Pushing the frontiers of audio generation

#104

Earlier quoted context omitted.

That's wild. Thanks for the links. Infinite AI Muzak I guess? I suppose it was inevitable.

Seems to be genuinely popular. I can see why, but when there's so much 'real music' out there, why not just listen to that and enrich yourself instead of bathing in fake nonsense? If you want jazz, just put on a jazz album — hell, even Kind of Blue will do. I'm not sure all (or any) of it actually is AI though. I assume that's coming very soon, but I suspect this stuff is cynically and methodically hand-composed. By…

Not longer than Steve Reich, John Cage, Philip Glass, or Ann Southam.

Nor for that matter Mozart, who wrote simple algorithmic compositions powered by dice. These were common musical games in his day.

Re: Pushing the frontiers of audio generation

#105
post #9

Is there a free (ad supported?) online tool without login that reads text that you paste into it? I often would like to listen to a blog post instead of reading it, but haven't found an easy, quick solution yet. I tried piping text through OpenAI's tts-1-hd, model and it is the first one I ever found that is human like enough for me to like listening to it. So I could write a tool for my own usecase that pipes the te…

I use ms edge for this exact use case. Works well enough on any platform

Yup, I second this. The only thing I use Edge for.

I did a bit of research and it seems to be, by far, the highest-quality TTS engine that is free and you can do things like pause and continue.

There are other options that have higher-quality voices, but they aren't free.

Re: Pushing the frontiers of audio generation

#106

Earlier quoted context omitted.

Seems to be genuinely popular. I can see why, but when there's so much 'real music' out there, why not just listen to that and enrich yourself instead of bathing in fake nonsense? If you want jazz, just put on a jazz album — hell, even Kind of Blue will do. I'm not sure all (or any) of it actually is AI though. I assume that's coming very soon, but I suspect this stuff is cynically and methodically hand-composed. By…

Not longer than Steve Reich, John Cage, Philip Glass, or Ann Southam. Nor for that matter Mozart, who wrote simple algorithmic compositions powered by dice. These were common musical games in his day.

You’re probably right, but I meant generative in the sense of building a machine that does the entire process.

Re: Pushing the frontiers of audio generation

#107
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

I got a similar feeling. I think it was overdoing the ums and uhhs for something trying to sound like an even slightly professional podcast kind of sound.

Re: Pushing the frontiers of audio generation

#108
This is a "holy shit" moment for me, and I consider myself fairly jaded. If you listen closely you can tell it's a little off, but about halfway through I could clearly feel my brain click into a different mode where it believed what it was hearing was real.

Re: Pushing the frontiers of audio generation

#109
post #43

But what's the end goal and audience here? I don't believe people will resonate with robots making "um" and "ohs" because people usually resonate with an artist, a producer, a writer, a singer etc. A human layer with which people can empathize is essential. This can work as long as people are deceived and don't know there is no human behind it. If however i find out that a video is AI -generated i instantly lose inte…

It’s amazing for reading articles

Re: Pushing the frontiers of audio generation

#110

Earlier quoted context omitted.

I use ms edge for this exact use case. Works well enough on any platform

Yup, I second this. The only thing I use Edge for. I did a bit of research and it seems to be, by far, the highest-quality TTS engine that is free and you can do things like pause and continue. There are other options that have higher-quality voices, but they aren't free.

I use the wavenet extension and use my 1M character free quota from Google cloud
Post reply on HN