But what's the end goal and audience here? I don't believe people will resonate with robots making "um" and "ohs" because people usually resonate with an artist, a producer, a writer, a singer etc. A human layer with which people can empathize is essential. This can work as long as people are deceived and don't know there is no human behind it. If however i find out that a video is AI -generated i instantly lose inte…
Pushing the frontiers of audio generation
101–110 of 114 posts
Re: Pushing the frontiers of audio generation
#102But what's the end goal and audience here? I don't believe people will resonate with robots making "um" and "ohs" because people usually resonate with an artist, a producer, a writer, a singer etc. A human layer with which people can empathize is essential. This can work as long as people are deceived and don't know there is no human behind it. If however i find out that a video is AI -generated i instantly lose inte…
I think they absolutely will, because "resonating" is not a material phenomenon, it's something people decide that they're doing. Your connection with an actor on television is not an actual connection. Most of acting is learning the times and length to be silent while making a particular face (dictated by the director) in order for the audience to project feelings and thoughts onto you. You're thinking about your camera blocking, or your groceries, and your audience sees you thinking about some plot point in a fictional world.
I've got a theory that we severely damaged a generation of girls by inundating them with images of girls their own age singing songs and acting parts all written and directed by middle-aged men - ones who chose as a profession to write songs in the voices of, write fiction in the voices of, and to direct, photograph and choreograph in person, tween girls. Their models of themselves have come from looking at these depictions of girls, who were never allowed to speak for themselves, and resonating.
Re: Pushing the frontiers of audio generation
#103Earlier quoted context omitted.
We're used to hearing some kind of identity behind voices -- we unconsciously sense clusters of vocabulary, intonation patterns, ticks, frequent interruption vs quiet patience, silence tolerance, response patterns to various triggers, etc that communicate a coherent person of some kind. We may not know that a given speaker is a GenX Methodist from Wisconsin that grew up at skate parks in the suburbs, but we hear clus…
To be fair, the majority of podcasts are from a group of generic white guys, and they almost sound identical to these AI generated ones. The AI actually seems to to do a better job too.
When I listened to the audio samples before coming to the comments, I thought: "oh, like those totally lifeless and bland U.S. accents from podcasts, YT, etc."
I wouldn't associate it with skin colour or gender though at all. I've no idea why you'd go there - any skin colour and any gender is absolutely welcomed into the fold of U.S. cultural production, if they can produce bland generic "content" sincerely enough, it seems to me.
Disclaimer: many U.S. accents are interesting and wonderful (Colorado; Tom Waits), they don't all sound generic and bland. I have U.S. friends therefore I can pass judgment (TM).
Re: Pushing the frontiers of audio generation
#104Earlier quoted context omitted.
That's wild. Thanks for the links. Infinite AI Muzak I guess? I suppose it was inevitable.
Seems to be genuinely popular. I can see why, but when there's so much 'real music' out there, why not just listen to that and enrich yourself instead of bathing in fake nonsense? If you want jazz, just put on a jazz album — hell, even Kind of Blue will do. I'm not sure all (or any) of it actually is AI though. I assume that's coming very soon, but I suspect this stuff is cynically and methodically hand-composed. By…
Nor for that matter Mozart, who wrote simple algorithmic compositions powered by dice. These were common musical games in his day.
Re: Pushing the frontiers of audio generation
#105Is there a free (ad supported?) online tool without login that reads text that you paste into it? I often would like to listen to a blog post instead of reading it, but haven't found an easy, quick solution yet. I tried piping text through OpenAI's tts-1-hd, model and it is the first one I ever found that is human like enough for me to like listening to it. So I could write a tool for my own usecase that pipes the te…
I use ms edge for this exact use case. Works well enough on any platform
I did a bit of research and it seems to be, by far, the highest-quality TTS engine that is free and you can do things like pause and continue.
There are other options that have higher-quality voices, but they aren't free.
Re: Pushing the frontiers of audio generation
#106Earlier quoted context omitted.
Seems to be genuinely popular. I can see why, but when there's so much 'real music' out there, why not just listen to that and enrich yourself instead of bathing in fake nonsense? If you want jazz, just put on a jazz album — hell, even Kind of Blue will do. I'm not sure all (or any) of it actually is AI though. I assume that's coming very soon, but I suspect this stuff is cynically and methodically hand-composed. By…
Not longer than Steve Reich, John Cage, Philip Glass, or Ann Southam. Nor for that matter Mozart, who wrote simple algorithmic compositions powered by dice. These were common musical games in his day.
Re: Pushing the frontiers of audio generation
#107While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.
Re: Pushing the frontiers of audio generation
#108Re: Pushing the frontiers of audio generation
#109But what's the end goal and audience here? I don't believe people will resonate with robots making "um" and "ohs" because people usually resonate with an artist, a producer, a writer, a singer etc. A human layer with which people can empathize is essential. This can work as long as people are deceived and don't know there is no human behind it. If however i find out that a video is AI -generated i instantly lose inte…
Re: Pushing the frontiers of audio generation
#110Earlier quoted context omitted.
I use ms edge for this exact use case. Works well enough on any platform
Yup, I second this. The only thing I use Edge for. I did a bit of research and it seems to be, by far, the highest-quality TTS engine that is free and you can do things like pause and continue. There are other options that have higher-quality voices, but they aren't free.