Live data from Hacker News

Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

stability.ai

41–50 of 210 posts

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#41
post #30
post #6

The bluegrass one is super weird. I can’t identify exactly why.

I think the super weird part is that it's not great? I understand this is most likely very impressive technologically but musically it is disjointed, inconsistent and fake sounding. Most of the "music" examples have weird phrasing and confusing harmonic rhythm. Kudos to stability.ai for achieving this as I am sure it took a lot of effort and this is a huge leap forward in terms of generation of audio by generative AI…

I feel like music composition is a fundamentally hard task for AI. Music production seems like it should be a lot easier but I haven’t seen that

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#42

It's interesting tech but none of the musical pieces impressed me (I play multiple instruments and have written and arranged music), most sounded too repetitive and not very imaginative. This is also an issue with diffusion based art AI in general, its good at a limited set of things but gets rather repetitive after a while. I could see using this as background music where quality is not important, like in games, tho…

I'm surprised that you even leap to people like Hans Zimmer and others.

The people we need to worry about are aallll of the people earning a living for everything else like background music for Indi games, Ambient music etc.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#43
The solo piano was interesting because of how clean it is. I can imagine going from that sample to a score without too much difficulty. Once it's in a symbolic format it becomes much more flexible and re-usable.

While this does not seem to be the trend I hope more gen ai in the audio and visual realms start to produce more structured / symbolic output. For example, if I were Adobe I would be training models, not to output full images, but either layers or brush strokes and tool pallet usage. Same for organizations that have all the component tracks of music to work with.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#44

It's interesting tech but none of the musical pieces impressed me (I play multiple instruments and have written and arranged music), most sounded too repetitive and not very imaginative. This is also an issue with diffusion based art AI in general, its good at a limited set of things but gets rather repetitive after a while. I could see using this as background music where quality is not important, like in games, tho…

I thought the, "epic trailer music intense tribal percussion and brass" was pretty good. Rather, good enough for something like a video game where the game engine is dynamically generating music based on the present situation in the game.

I could easily find that music entertaining if it started playing the moment my character triggered a trap and suddenly, "the floor is lava" or my character enters a scene with the quest of winning over one of the love interests =)

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#45

Earlier quoted context omitted.

How is Spotify for finding new music based on your tastes? I’ve only used Amazon and Pandora; Amazon is quite poor, Pandora is pretty good. I suspect (although, without proof) that if a service can’t suggest new music, it will have trouble generating new music as well. Anyway, I very much would rather run this sort of thing locally. You could just manually set your taste profile. Plus, music can be quite personal, im…

> How is Spotify for finding new music based on your tastes? I haven't tried any alternatives really, but so far for me I'd say decent. I put on an album, and once over it'll play similarish stuff. If I don't like a song I'll skip it and it seems to incorporate that feedback. Only thing is that it doesn't seem to be too adventurous and it adheres rather strictly to the local context. Meaning, if I played a stoner roc…

I'll +1. I'm generally a FOSS guy, so if Spotify hadn't helped so much in discovering new music that I really enjoy, I'd be potentially acquiring music via questionable means and just playing them as audio files directly.

There are two companies that have done well by Gabe Newell's "piracy is a service problem - not a price problem" position: Valve/Steam (who also contribute to FOSS through Proton and SteamOS which I heavily appreciate), and Spotify. Spotify makes discovering and aggregating music so easy that the alternatives don't seem appealing.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#47
post #21

Earlier quoted context omitted.

This is not the way.

Why not? Mostly for private use in my case. SDXL has created some beautiful works of art in my experiments and I would love to have a similar experience in the music world.

Do you know where the sounds came from that you like so much? I think such a perspective is only reachable if you do not.

I recommend learning about that before deciding it’s satisfactory to reduce it to an algorithm suited for copying.

It’ll enrich your life. Endless copies will not. They take that music, that emergence of order out of chaos, and return it back to chaos.

It’s void.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#49
post #32
post #4

I keep thinking back to when we didn't have stabilityai and it was just google and meta teasing us with mouth watering papers but never letting us touch them. I'm so thankful stability exists.

Stability is great but Meta's MusicGen is available with code and weights while this isn't so that's a really odd place to make that comparison and complaint.

Before stable diffusion, nobody released weights at all. Meta et al only started sharing their models with the world when they realized how fast a developer ecosystem was building around the best models.

Without stability, all of AI would still be closed and opaque.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#50
Does this model support / "understand" concepts of spatial audio? For example, something like "an alarm moving around you in a circle".

When AudioGen was announced this was my first question, but from what I've been able to test the model just ignores spatial audio prompts.

Unfortunately I haven't been able to find any discussion or interest in online discussion about the importance / significance of spatial audio. Why not?

Post reply on HN