While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.
Pushing the frontiers of audio generation
51–60 of 114 posts
Re: Pushing the frontiers of audio generation
#52Re: Pushing the frontiers of audio generation
#53Earlier quoted context omitted.
Totally agree. Maybe it’s just the clips they chose, but it feels overfit on the weird conversational elements that make it impressive? Like the “oh yeahs” from the other person when someone is speaking. It is cool to see that natural flow in a conversation generated by a model, but there’s waaaay too much of it in these examples to sound natural. And I say all that completely slackjawed that this is possible.
> Like the “oh yeahs” from the other person when someone is speaking. I bet that if you select a British accent you will get fewer of them.
Hmm.... Scottish, Welsh, Irish (Nor'n) or English? If English, North or South? If North, which city? Brummie? Scouse? If South, London? Cockney or Multicultural London English [0]?
[0] https://en.wikipedia.org/wiki/Multicultural_London_English
Re: Pushing the frontiers of audio generation
#54Earlier quoted context omitted.
To be fair, the majority of podcasts are from a group of generic white guys, and they almost sound identical to these AI generated ones. The AI actually seems to to do a better job too.
Citation absolutely needed. You call this fair? > the majority of podcasts are from a group of generic white guys
Re: Pushing the frontiers of audio generation
#55Earlier quoted context omitted.
> Like the “oh yeahs” from the other person when someone is speaking. I bet that if you select a British accent you will get fewer of them.
> a British accent Hmm.... Scottish, Welsh, Irish (Nor'n) or English? If English, North or South? If North, which city? Brummie? Scouse? If South, London? Cockney or Multicultural London English [0]? [0] https://en.wikipedia.org/wiki/Multicultural_London_English
Re: Pushing the frontiers of audio generation
#56Earlier quoted context omitted.
> Like the “oh yeahs” from the other person when someone is speaking. I bet that if you select a British accent you will get fewer of them.
> a British accent Hmm.... Scottish, Welsh, Irish (Nor'n) or English? If English, North or South? If North, which city? Brummie? Scouse? If South, London? Cockney or Multicultural London English [0]? [0] https://en.wikipedia.org/wiki/Multicultural_London_English
Castlebridge is 10 minutes away by car. Madness!
Re: Pushing the frontiers of audio generation
#57YouTube videos are already infested with insufferable AI elevator background "music". Even some channels that were previously good are using it. On the bright side, you can stop watching these channels and have more time for serious things.
> AI elevator background "music". What are some examples? I haven't encountered this.
Almost all of the results will not consist of 'jazz' in any real sense, but instead a collection of uncanny melodies and chord progressions that wonder around going nowhere, traditionally accompanied by an obscenely eye-offending diffusion model-generated mishmash of seasonal tropes and incongruent interior design choices. Often, it's MIDI bossa nova presumably written by either a machine or someone who's only ever heard a few bars of music at a time and has no idea that 'feel' or 'soul' are a thing.
Re: Pushing the frontiers of audio generation
#58The voices are impressive (I can't tell the difference as a non native speaker) but their "personality" sounds extremely annoying lmao
Re: Pushing the frontiers of audio generation
#59Is there a free (ad supported?) online tool without login that reads text that you paste into it? I often would like to listen to a blog post instead of reading it, but haven't found an easy, quick solution yet. I tried piping text through OpenAI's tts-1-hd, model and it is the first one I ever found that is human like enough for me to like listening to it. So I could write a tool for my own usecase that pipes the te…
Re: Pushing the frontiers of audio generation
#60Earlier quoted context omitted.
Whether this stops at the uncanny valley or progresses to specific "AI celebrity" voices, I'm left thinking the engineers involved in this never stopped to think carefully about whether this ought to be done in the first place.
"Surely my genAI product won't be used to spam zero-effort slop all over the internet!" - guy whose genAI product will definitely be used to spam zero-effort slop all over the internet.
We are witnessing in real time the answer to why 'The Matrix' was set when it was. Once AI takes over there is no future culture.