Live data from Hacker News

Pushing the frontiers of audio generation

deepmind.google

91–100 of 114 posts

Re: Pushing the frontiers of audio generation

#91
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

It doesn't feel any different to me than listening to a random radio station where I don't know who is speaking. I didn't feel any uncanny valley but I'm not an English native speaker so I might miss some nuances. However there are relatively few English native speakers around the world so this might not be a problem for us.

The problem is that people talking over each other is not a format I long to listen to.

Re: Pushing the frontiers of audio generation

#92
post #82

Earlier quoted context omitted.

I think their main target is corporate creative jobs. Background music to ads/videos/etc. And just like with all AI, they will eat the jobs that support the rest of the system, making it a one and done. It will give a one time boost, and then be stuck at that level because creatives won't have the jobs that allowed them to add to the domain. In this case new music styles. New techniques. It's literally eating the see…

this is very spot on. There are tons of artists who have a job so they can sustain their own personal creativity.

Almost everyone has a job to sustain their actual interests. Some of them happen to be musicians, writers, etc. Others play football, go fishing, talk to friends. There is nothing special in there. All of us will keep doing what we like to do even after AIs become the tool of mainstream creativity.

Re: Pushing the frontiers of audio generation

#93
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

For me it isn’t uncanny from a lack of humanity. Rather, it triggers all my “fake and shallow” personality biases. It certainly sounds human enough, just not the type of humans I like.

Re: Pushing the frontiers of audio generation

#94
post #92
post #82

Earlier quoted context omitted.

this is very spot on. There are tons of artists who have a job so they can sustain their own personal creativity.

Almost everyone has a job to sustain their actual interests. Some of them happen to be musicians, writers, etc. Others play football, go fishing, talk to friends. There is nothing special in there. All of us will keep doing what we like to do even after AIs become the tool of mainstream creativity.

Yes but many people have jobs in their fields, which benefits their personal and/or public creative endeavours.

I’ve learnt things and been exposed to ideas developing software for work that I simply wouldn’t have if I was only doing it in my spare time.

Re: Pushing the frontiers of audio generation

#95
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

There’s a certain fakeness to the rhythm of the space between words. Particularly the “uh” and “um” filler sounds. To me it sounds like they always either come in abnormally early or late after speaking those sounds

Re: Pushing the frontiers of audio generation

#96
post #43

But what's the end goal and audience here? I don't believe people will resonate with robots making "um" and "ohs" because people usually resonate with an artist, a producer, a writer, a singer etc. A human layer with which people can empathize is essential. This can work as long as people are deceived and don't know there is no human behind it. If however i find out that a video is AI -generated i instantly lose inte…

> what's the end goal and audience here?

1. Voice acting for low-budget/no-budget animations and games.

2. Billions of youtube "top 50 building demolitions" where the forgettable presentation is narrated by forgettable AI. Now we'll get "podcast style" conversation narration over those videos. Instead of bailing after 30 sec with regret, you might make a whole minute.

3. Reaction videos? Sometimes I weaken. I want to see a random person's reaction to their "first time listening" to the famous song they somehow have never heard until this moment. If we humans lower ourselves to reaction videos, we'll watch/listen to AI chatting to itself about things we love. Once the content gets "spicy", beyond the potato salad google demos, the floodgates will open. God help us.

Re: Pushing the frontiers of audio generation

#97
post #81

We've been using this at work to get inside of our customer's perspective. It's helpful to throw eg a bunch of point-of-sale data sync challenges into Notebook LM and eg pass a 10 minute audio to the team so they can understand where our work fits in.

I’ve cut and pasted weeks of Slack conversations into NotebookLM and it was quite entertaining to then listen to a Podcast talking humorously about all the arguments in the #management channel.

For someone who has never used NotebookLM, how would I get started doing this?

Re: Pushing the frontiers of audio generation

#98
post #71

Earlier quoted context omitted.

I think their main target is corporate creative jobs. Background music to ads/videos/etc. And just like with all AI, they will eat the jobs that support the rest of the system, making it a one and done. It will give a one time boost, and then be stuck at that level because creatives won't have the jobs that allowed them to add to the domain. In this case new music styles. New techniques. It's literally eating the see…

Assuming you are right and that we will miss a generation of creatives and AI keeps making crap, why can't the creative field regrow. AI won't remove creativity from human genes. As people get fed up with AI generated crap, companies will start to pay very good money to the few remaining good human creatives in order to differentiate themselves. The field will then be seen as desirable, people will start working hard…

It makes sense to me.

Step 0. Some People make novel art like a jingle that is unlike anything yet.

Step 1. Early use of said jingle creates a buzz and generated good sales results.

Step 2. It gets copied everywhere and by everyone. It is now a meme.

This is the step I think where generative AI can help. Slightly transform existing art to fit a particular purpose. This lets businesses save money by not paying humans do this work.

Problem is we don't know where the next person or when this step 0 comes from... When we soak up all the "slack" and send all the "money" to the top because lets face it that's how it will work. The money "saved" from AI won't make goods and services cheaper by any significant measure. We will still have to pay as much as we can afford to pay.

Re: Pushing the frontiers of audio generation

#99

Earlier quoted context omitted.

I think their main target is corporate creative jobs. Background music to ads/videos/etc. And just like with all AI, they will eat the jobs that support the rest of the system, making it a one and done. It will give a one time boost, and then be stuck at that level because creatives won't have the jobs that allowed them to add to the domain. In this case new music styles. New techniques. It's literally eating the see…

> It's literally eating the seed corn where the sprouts are the creatives working in the boring commercial jobs that allow them to practice/become experts in the tools/etc that they then build up it all. This is a big problem that needs to be talked about more, the endgoal of AI seems to be quite grim for jobs and generally for humans. Where will this pure profit lead to? If all advertising will be generated who will…

Reminds me of that famous clip from mad men where don suddenly realizes that if lucky strike can't say its cigarettes are safe, neither can its competitors and came up with "it's toasted".

In general, I have a feeling double digit growth forever is impossible. Facebook and Google both reported YoY growth in 15%+ this week iirc and I have a feeling they are only able to achieve this by destroying either competitors or adjacent industries rather than by "making the pie bigger". It will end at some point.

Re: Pushing the frontiers of audio generation

#100

Earlier quoted context omitted.

We're used to hearing some kind of identity behind voices -- we unconsciously sense clusters of vocabulary, intonation patterns, ticks, frequent interruption vs quiet patience, silence tolerance, response patterns to various triggers, etc that communicate a coherent person of some kind. We may not know that a given speaker is a GenX Methodist from Wisconsin that grew up at skate parks in the suburbs, but we hear clus…

Agreed. To me it sounds like bad voice-over actors reading from a script. So the natural parts of a conversation where you might say the wrong thing and step back to correct yourself are all gone. Impressive for sure.

Yup. Plus the interactions you'd expect for instance in terms of matching style of voice in a normal discussion are missing. That being said it still sounds pretty impressive.
Post reply on HN