Live data from Hacker News

Pushing the frontiers of audio generation

deepmind.google

71–80 of 114 posts

Re: Pushing the frontiers of audio generation

#71
post #50

Earlier quoted context omitted.

"Surely my genAI product won't be used to spam zero-effort slop all over the internet!" - guy whose genAI product will definitely be used to spam zero-effort slop all over the internet.

I think their main target is corporate creative jobs. Background music to ads/videos/etc. And just like with all AI, they will eat the jobs that support the rest of the system, making it a one and done. It will give a one time boost, and then be stuck at that level because creatives won't have the jobs that allowed them to add to the domain. In this case new music styles. New techniques. It's literally eating the see…

Assuming you are right and that we will miss a generation of creatives and AI keeps making crap, why can't the creative field regrow. AI won't remove creativity from human genes.

As people get fed up with AI generated crap, companies will start to pay very good money to the few remaining good human creatives in order to differentiate themselves. The field will then be seen as desirable, people will start working hard for to get these jobs, companies will take apprentices hoping they will become masters later, etc... We may lose a generation, but certainly not the entire future.

Of course, it is just one of many possible futures, but I think the most likely if you take your assumptions as a postulate. It may turn out that AIs end up not displacing creative jobs too much, or going the other way, that AIs end up being truly creative, building their own culture together with humans, or not.

Re: Pushing the frontiers of audio generation

#72
post #16

Earlier quoted context omitted.

> Like the “oh yeahs” from the other person when someone is speaking. I bet that if you select a British accent you will get fewer of them.

> a British accent Hmm.... Scottish, Welsh, Irish (Nor'n) or English? If English, North or South? If North, which city? Brummie? Scouse? If South, London? Cockney or Multicultural London English [0]? [0] https://en.wikipedia.org/wiki/Multicultural_London_English

When people outside the British isles (esp. Americans) say "British accent", they almost invariably mean (British) English, and usually the "received pronunciation" accent that British media generally uses.

They do not mean Irish or Scottish accents; if they did, they would have said exactly that, because those accents are quite different from standard (British) English accents. So different, in fact, that even Americans can readily tell the difference, when they frequently have some trouble telling English and Australian accents apart.

Also, to most English speakers, "English accent" doesn't make much sense, because "English" is the language. It sounds like saying a German speaker, speaking German, has a "German accent". Saying "British accent" differentiates the language (English, spoken by people worldwide) from the accent (which refers to one part of one country that uses that language).

Re: Pushing the frontiers of audio generation

#73

Earlier quoted context omitted.

Can you post a link to some like that? Because when I search "jazz" on YT I'm just getting legit music videos and jazz playlists -- stuff like Norah Jones, top 100 jazz classics playlists, etc. But I assume that search results are personalized.

Lucky you! Sure. I just tried in private browsing mode, and got mostly the same. Here are a few of the very first results I get for 'jazz': https://www.youtube.com/watch?v=xhL3Cb740VY https://www.youtube.com/watch?v=8UXFapv_kFI https://www.youtube.com/watch?v=nKNnzbi-v9E https://www.youtube.com/watch?v=ABmQvH5K75w https://www.youtube.com/watch?v=-jgEswq9ZlI Some are worse than others.

That's wild. Thanks for the links. Infinite AI Muzak I guess? I suppose it was inevitable.

Re: Pushing the frontiers of audio generation

#74
post #20

Earlier quoted context omitted.

Totally agree. Maybe it’s just the clips they chose, but it feels overfit on the weird conversational elements that make it impressive? Like the “oh yeahs” from the other person when someone is speaking. It is cool to see that natural flow in a conversation generated by a model, but there’s waaaay too much of it in these examples to sound natural. And I say all that completely slackjawed that this is possible.

I love the technology, but I really don't want AI to sound like this. Imagine being stuck on a call with this. > "Hey, so like, is there anything I can help you with today?" > "Talk to a person." > "Oh wow, right. (chuckle) You got it. Well, before I connect you, can you maybe tell me a little bit more about what problem you're having? For example, maybe it's something to do with..."

Reminds me of the robots from the Sirius cybernetics corporation. “Your plastic pal who’s fun to be with.”

Re: Pushing the frontiers of audio generation

#76
post #9

Is there a free (ad supported?) online tool without login that reads text that you paste into it? I often would like to listen to a blog post instead of reading it, but haven't found an easy, quick solution yet. I tried piping text through OpenAI's tts-1-hd, model and it is the first one I ever found that is human like enough for me to like listening to it. So I could write a tool for my own usecase that pipes the te…

firefox does this directly. Reader mode has a headphones symbol to read webpage text.

Re: Pushing the frontiers of audio generation

#77

Earlier quoted context omitted.

Lucky you! Sure. I just tried in private browsing mode, and got mostly the same. Here are a few of the very first results I get for 'jazz': https://www.youtube.com/watch?v=xhL3Cb740VY https://www.youtube.com/watch?v=8UXFapv_kFI https://www.youtube.com/watch?v=nKNnzbi-v9E https://www.youtube.com/watch?v=ABmQvH5K75w https://www.youtube.com/watch?v=-jgEswq9ZlI Some are worse than others.

That's wild. Thanks for the links. Infinite AI Muzak I guess? I suppose it was inevitable.

Seems to be genuinely popular. I can see why, but when there's so much 'real music' out there, why not just listen to that and enrich yourself instead of bathing in fake nonsense? If you want jazz, just put on a jazz album — hell, even Kind of Blue will do.

I'm not sure all (or any) of it actually is AI though. I assume that's coming very soon, but I suspect this stuff is cynically and methodically hand-composed.

By the way: I have nothing against generative composition! Brian Eno has been doing this stuff longer than anyone else, and it's very cool. I'm sure you could make some 'generative jazz' that's actually distinctive and artistic, but this isn't it.

Re: Pushing the frontiers of audio generation

#78
post #29
post #4

While it is impressive and I like to follow the advancements in this field, it is incredibly frustrating to listen to. I can't put my finger on why exactly. It's definitely closer to human-sounding, but the uncanny valley is so deep here that I find myself thinking "I just want the point, not the fake personality that is coming with it". I can't make it through a 30s demo.

they all sound like valley-people, complete with the raspy voice and everything

[deleted]

Re: Pushing the frontiers of audio generation

#79
post #32

Earlier quoted context omitted.

It's like their training set was made up entirely of awkward podcaster banter.

At least 83% Leo Laporte.

If I turn the volume down to the point that I only hear the cadence/rhythm of the voices, but can no longer make out the words, it sounds like any, “This Week in…” podcast.

Re: Pushing the frontiers of audio generation

#80
post #23

It looks like lately a lot of progress have been made in audio generation / audio understanding (everything related to speech, I mean). Is this related to LLM, or is this a completely different branch of AI, and is it just a coincidence? I am curious.

It's very related to LLMs. Though instead of text tokens you are working with audio tokens (e.g. from SoundStream). Then you go to audio corpus, instead of text corpus.
Post reply on HN