Live data from Hacker News

Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

stability.ai

21–30 of 210 posts

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#21
post #9

This is gonna be great to finetune on. There's only so many boards of canada/aphex twin songs out there but I wish there were more and this will let us generate more.

This is not the way.

Why not? Mostly for private use in my case. SDXL has created some beautiful works of art in my experiments and I would love to have a similar experience in the music world.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#22
post #21

Earlier quoted context omitted.

This is not the way.

Why not? Mostly for private use in my case. SDXL has created some beautiful works of art in my experiments and I would love to have a similar experience in the music world.

Come on, be creative and make something new instead of copying someone else.

It’s just kind of lame imo.

“Mostly” private use? Mmm. :thumbs down emoji:

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#24
post #21

Earlier quoted context omitted.

Why not? Mostly for private use in my case. SDXL has created some beautiful works of art in my experiments and I would love to have a similar experience in the music world.

Come on, be creative and make something new instead of copying someone else. It’s just kind of lame imo. “Mostly” private use? Mmm. :thumbs down emoji:

I meant private use and maybe share with a few friends. I actually agree with you that we probably shouldn't finetune on great artists and try to sell the output without modification or added creativity. Private or close friends sharing is fun and life enriching and inspiring though in my eyes.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#25
post #3

Thank you for sharing! On a tangent: I'm wondering if there are any good open source models/libraries to reconstruct audio quality. I'm thinking about an end-to-end open source alternative to something like Adobe Podcast [1] to make noisy recordings sound professional. Anecdotally it's supposed to be very good. In a recent search, I haven't found anything convincing. In my naive view this tasks seems much simpler tha…

There seems to have been a fork in the road: On one side the tech for literal denoising has stagnated a bit. It’s a very hard problem to remove all noise while keeping things like transients. On the other side, AI is being rapidly developed for it’s ability to denoise by recreating the recording, just without the noise.

In our denoiser (see other comment), we worked on combining these two forks. That’s how we can mathematically guarantee great audio quality.

This combination was non-trivial as training old school DSP denoisers is not easily possible. We’ll describe the math needed in our paper. We hope our publication will help the wider community work not just on denoising but also tasks like automatic mixing.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#26
post #7

Everything looks very convincing apart from the airpane pilot and the sound effects. They sound very weird as if one is hallucinating

The airplane one just sounds like a foreign language over a bad intercom, I think that could still be useful for some stuff.

Perhaps because generating good white noise requires randomness without autocorrelation or detectable patterns.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#27

Now imagine Spotify using this to generate individual earworms for everybody based on their personal tastes (likes, playlists). Yes, AI is partly hype, but had someone told me this even two years ago, I wouldn't have believed it.

This is why it’s vital that AI is openly available. Imagine a world where Spotify is the only company that can do that, and they use it to make sure they never pay royalties again.

How is Spotify for finding new music based on your tastes? I’ve only used Amazon and Pandora; Amazon is quite poor, Pandora is pretty good. I suspect (although, without proof) that if a service can’t suggest new music, it will have trouble generating new music as well.

Anyway, I very much would rather run this sort of thing locally. You could just manually set your taste profile. Plus, music can be quite personal, imagine you start listening to too much music inspired by The Cure and suddenly Amazon starts advertising black makeup and antidepressants or something like that, it would be too disconcerting.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#28

Earlier quoted context omitted.

This is why it’s vital that AI is openly available. Imagine a world where Spotify is the only company that can do that, and they use it to make sure they never pay royalties again.

How is Spotify for finding new music based on your tastes? I’ve only used Amazon and Pandora; Amazon is quite poor, Pandora is pretty good. I suspect (although, without proof) that if a service can’t suggest new music, it will have trouble generating new music as well. Anyway, I very much would rather run this sort of thing locally. You could just manually set your taste profile. Plus, music can be quite personal, im…

> How is Spotify for finding new music based on your tastes?

I haven't tried any alternatives really, but so far for me I'd say decent. I put on an album, and once over it'll play similarish stuff. If I don't like a song I'll skip it and it seems to incorporate that feedback.

Only thing is that it doesn't seem to be too adventurous and it adheres rather strictly to the local context. Meaning, if I played a stoner rock track, it'll continue suggesting stoner rock and not much else, even though I have quite varied music favorited in my library.

Overall though I've found a lot of new bands I enjoy that way so, positive experience for me.

edit: as an example, here are the two most recent ones it suggested where I ended up buying the albums on Bandcamp. Both have quite few monthly listeners (1-2k), so not what I'd call mainstream.

SUIR https://open.spotify.com/artist/6zOeQ2hyNfqi9UMHtyTSlF

Mount Hush https://open.spotify.com/artist/13clfeXxTPsDsqzSlLIBZJ

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#30
post #6

The bluegrass one is super weird. I can’t identify exactly why.

I think the super weird part is that it's not great? I understand this is most likely very impressive technologically but musically it is disjointed, inconsistent and fake sounding. Most of the "music" examples have weird phrasing and confusing harmonic rhythm.

Kudos to stability.ai for achieving this as I am sure it took a lot of effort and this is a huge leap forward in terms of generation of audio by generative AI.

However as a musician (BMus and MMus at 2 different conservatoires) I think it's important to say that the job risk being experienced by creative writers will not be extending to musicians... yet.

Post reply on HN