Live data from Hacker News

Pushing the frontiers of audio generation

deepmind.google

61–70 of 114 posts

Re: Pushing the frontiers of audio generation

#61
post #56

Earlier quoted context omitted.

> a British accent Hmm.... Scottish, Welsh, Irish (Nor'n) or English? If English, North or South? If North, which city? Brummie? Scouse? If South, London? Cockney or Multicultural London English [0]? [0] https://en.wikipedia.org/wiki/Multicultural_London_English

Need to increase your granularity a bit. I live in Wexford Town, Ireland, and the other day I was chatting to a person that told me their old schoolmates from Castlebridge are making fun of their accent changing since moving from their hometown. Castlebridge is 10 minutes away by car. Madness!

Yeah, totally agree. Here's a useful link for non-Brits, that goes into a bit more detail:

https://accentbiasbritain.org/accents-in-britain/

Also, we have yet to define precisely define what is meant by 'British'. This probably needs a "20 falsehoods people believe about..."-type article.

Re: Pushing the frontiers of audio generation

#63

Earlier quoted context omitted.

We're used to hearing some kind of identity behind voices -- we unconsciously sense clusters of vocabulary, intonation patterns, ticks, frequent interruption vs quiet patience, silence tolerance, response patterns to various triggers, etc that communicate a coherent person of some kind. We may not know that a given speaker is a GenX Methodist from Wisconsin that grew up at skate parks in the suburbs, but we hear clus…

Agreed. To me it sounds like bad voice-over actors reading from a script. So the natural parts of a conversation where you might say the wrong thing and step back to correct yourself are all gone. Impressive for sure.

every step of technological advancement builds on top of the previous one.

now it's bad voice actors, in 2 years it'll be great ones

Re: Pushing the frontiers of audio generation

#64

Earlier quoted context omitted.

Citation absolutely needed. You call this fair? > the majority of podcasts are from a group of generic white guys

https://podcastcharts.byspotify.com/ keep the Pareto distribution in mind

I did the best fast research I could given not wanting to spend more than 20 minutes on it and came to this result (aprox): - Mixed/Diverse: 48.0% - White Men: 35.0% - Women: 8.0% - Non-White: 6.0% - White Woman: 2.0% - Non-White Woman: 1.0%

Re: Pushing the frontiers of audio generation

#65
post #50

Earlier quoted context omitted.

"Surely my genAI product won't be used to spam zero-effort slop all over the internet!" - guy whose genAI product will definitely be used to spam zero-effort slop all over the internet.

I think their main target is corporate creative jobs. Background music to ads/videos/etc. And just like with all AI, they will eat the jobs that support the rest of the system, making it a one and done. It will give a one time boost, and then be stuck at that level because creatives won't have the jobs that allowed them to add to the domain. In this case new music styles. New techniques. It's literally eating the see…

> It's literally eating the seed corn where the sprouts are the creatives working in the boring commercial jobs that allow them to practice/become experts in the tools/etc that they then build up it all.

This is a big problem that needs to be talked about more, the endgoal of AI seems to be quite grim for jobs and generally for humans. Where will this pure profit lead to? If all advertising will be generated who will want to have anything to do with all the products they’re advertising?

Re: Pushing the frontiers of audio generation

#66
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

While it is impressive and I like to follow the advancements in this field...

Please don't think that I'm trying to suggest... anything . It's just that I'm getting used to read this pattern in the output of LLMs. "While this and that is great...". Maybe we're mimicking them now? I catch myself using these disclaimers even in spoken language.

Re: Pushing the frontiers of audio generation

#67
post #35

Earlier quoted context omitted.

> AI elevator background "music". What are some examples? I haven't encountered this.

Just search 'jazz' on YouTube. Almost all of the results will not consist of 'jazz' in any real sense, but instead a collection of uncanny melodies and chord progressions that wonder around going nowhere, traditionally accompanied by an obscenely eye-offending diffusion model-generated mishmash of seasonal tropes and incongruent interior design choices. Often, it's MIDI bossa nova presumably written by either a machi…

Can you post a link to some like that?

Because when I search "jazz" on YT I'm just getting legit music videos and jazz playlists -- stuff like Norah Jones, top 100 jazz classics playlists, etc.

But I assume that search results are personalized.

Re: Pushing the frontiers of audio generation

#68

Earlier quoted context omitted.

Just search 'jazz' on YouTube. Almost all of the results will not consist of 'jazz' in any real sense, but instead a collection of uncanny melodies and chord progressions that wonder around going nowhere, traditionally accompanied by an obscenely eye-offending diffusion model-generated mishmash of seasonal tropes and incongruent interior design choices. Often, it's MIDI bossa nova presumably written by either a machi…

Can you post a link to some like that? Because when I search "jazz" on YT I'm just getting legit music videos and jazz playlists -- stuff like Norah Jones, top 100 jazz classics playlists, etc. But I assume that search results are personalized.

Lucky you!

Sure. I just tried in private browsing mode, and got mostly the same. Here are a few of the very first results I get for 'jazz':

https://www.youtube.com/watch?v=xhL3Cb740VY

https://www.youtube.com/watch?v=8UXFapv_kFI

https://www.youtube.com/watch?v=nKNnzbi-v9E

https://www.youtube.com/watch?v=ABmQvH5K75w

https://www.youtube.com/watch?v=-jgEswq9ZlI

Some are worse than others.

Re: Pushing the frontiers of audio generation

#69
post #66
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

While it is impressive and I like to follow the advancements in this field... Please don't think that I'm trying to suggest... anything . It's just that I'm getting used to read this pattern in the output of LLMs. "While this and that is great...". Maybe we're mimicking them now? I catch myself using these disclaimers even in spoken language.

I like to preface negativity with a positive note. Maybe I am influenced in my word choice but my intent was to point out that this is a very, very impressive feat and I don't want to undermine it.
Post reply on HN