Live data from Hacker News

Pushing the frontiers of audio generation

deepmind.google

51–60 of 114 posts

Re: Pushing the frontiers of audio generation

#51
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

I get the feeling that this is useful for something that someone half-listens to.

Re: Pushing the frontiers of audio generation

#53
post #16

Earlier quoted context omitted.

Totally agree. Maybe it’s just the clips they chose, but it feels overfit on the weird conversational elements that make it impressive? Like the “oh yeahs” from the other person when someone is speaking. It is cool to see that natural flow in a conversation generated by a model, but there’s waaaay too much of it in these examples to sound natural. And I say all that completely slackjawed that this is possible.

> Like the “oh yeahs” from the other person when someone is speaking. I bet that if you select a British accent you will get fewer of them.

> a British accent

Hmm.... Scottish, Welsh, Irish (Nor'n) or English? If English, North or South? If North, which city? Brummie? Scouse? If South, London? Cockney or Multicultural London English [0]?

[0] https://en.wikipedia.org/wiki/Multicultural_London_English

Re: Pushing the frontiers of audio generation

#54

Earlier quoted context omitted.

To be fair, the majority of podcasts are from a group of generic white guys, and they almost sound identical to these AI generated ones. The AI actually seems to to do a better job too.

Citation absolutely needed. You call this fair? > the majority of podcasts are from a group of generic white guys

https://podcastcharts.byspotify.com/ keep the Pareto distribution in mind

Re: Pushing the frontiers of audio generation

#55
post #16

Earlier quoted context omitted.

> Like the “oh yeahs” from the other person when someone is speaking. I bet that if you select a British accent you will get fewer of them.

> a British accent Hmm.... Scottish, Welsh, Irish (Nor'n) or English? If English, North or South? If North, which city? Brummie? Scouse? If South, London? Cockney or Multicultural London English [0]? [0] https://en.wikipedia.org/wiki/Multicultural_London_English

[deleted]

Re: Pushing the frontiers of audio generation

#56
post #16

Earlier quoted context omitted.

> Like the “oh yeahs” from the other person when someone is speaking. I bet that if you select a British accent you will get fewer of them.

> a British accent Hmm.... Scottish, Welsh, Irish (Nor'n) or English? If English, North or South? If North, which city? Brummie? Scouse? If South, London? Cockney or Multicultural London English [0]? [0] https://en.wikipedia.org/wiki/Multicultural_London_English

Need to increase your granularity a bit. I live in Wexford Town, Ireland, and the other day I was chatting to a person that told me their old schoolmates from Castlebridge are making fun of their accent changing since moving from their hometown.

Castlebridge is 10 minutes away by car. Madness!

Re: Pushing the frontiers of audio generation

#57
post #35
post #25

YouTube videos are already infested with insufferable AI elevator background "music". Even some channels that were previously good are using it. On the bright side, you can stop watching these channels and have more time for serious things.

> AI elevator background "music". What are some examples? I haven't encountered this.

Just search 'jazz' on YouTube.

Almost all of the results will not consist of 'jazz' in any real sense, but instead a collection of uncanny melodies and chord progressions that wonder around going nowhere, traditionally accompanied by an obscenely eye-offending diffusion model-generated mishmash of seasonal tropes and incongruent interior design choices. Often, it's MIDI bossa nova presumably written by either a machine or someone who's only ever heard a few bars of music at a time and has no idea that 'feel' or 'soul' are a thing.

Re: Pushing the frontiers of audio generation

#59
post #9

Is there a free (ad supported?) online tool without login that reads text that you paste into it? I often would like to listen to a blog post instead of reading it, but haven't found an easy, quick solution yet. I tried piping text through OpenAI's tts-1-hd, model and it is the first one I ever found that is human like enough for me to like listening to it. So I could write a tool for my own usecase that pipes the te…

Good old Microsoft Sam? It'll sound like Stephen Hawking is reading it to you!

Re: Pushing the frontiers of audio generation

#60
post #50

Earlier quoted context omitted.

Whether this stops at the uncanny valley or progresses to specific "AI celebrity" voices, I'm left thinking the engineers involved in this never stopped to think carefully about whether this ought to be done in the first place.

"Surely my genAI product won't be used to spam zero-effort slop all over the internet!" - guy whose genAI product will definitely be used to spam zero-effort slop all over the internet.

I think their main target is corporate creative jobs. Background music to ads/videos/etc. And just like with all AI, they will eat the jobs that support the rest of the system, making it a one and done. It will give a one time boost, and then be stuck at that level because creatives won't have the jobs that allowed them to add to the domain. In this case new music styles. New techniques. It's literally eating the seed corn where the sprouts are the creatives working in the boring commercial jobs that allow them to practice/become experts in the tools/etc that they then build up it all. Their goal is cut the jobs that create their training data and the ecosystem that builds up/expands the domain. Everywhere AI touches will basically be 'stuck using Cobol' because AI will be frozen at the point in time where the energy infusing 'sprouts' all had their jobs replaced by AI and without them creating new output for AI to train on it's all ossified.

We are witnessing in real time the answer to why 'The Matrix' was set when it was. Once AI takes over there is no future culture.

Post reply on HN