Live data from Hacker News

Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

stability.ai

51–60 of 210 posts

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#52
post #32
post #4

I keep thinking back to when we didn't have stabilityai and it was just google and meta teasing us with mouth watering papers but never letting us touch them. I'm so thankful stability exists.

Stability is great but Meta's MusicGen is available with code and weights while this isn't so that's a really odd place to make that comparison and complaint.

Stability def helped push things forward , it even probably showed them that open source is inevitable

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#53

Earlier quoted context omitted.

This is why it’s vital that AI is openly available. Imagine a world where Spotify is the only company that can do that, and they use it to make sure they never pay royalties again.

How is Spotify for finding new music based on your tastes? I’ve only used Amazon and Pandora; Amazon is quite poor, Pandora is pretty good. I suspect (although, without proof) that if a service can’t suggest new music, it will have trouble generating new music as well. Anyway, I very much would rather run this sort of thing locally. You could just manually set your taste profile. Plus, music can be quite personal, im…

I've tried Amazon and Pandora. Spotify is so vastly superior at finding music I like it's in an entirely different league.

Having said that, starting about a year ago (maybe ~1.5 years?) Spotify started inserting obviously paid promotion tracks into my auto-generated "Daily Mix n" playlists and it seriously bothered me. When my playlists are made up of very specific genres of EDM and suddenly a pop song plays from a famous person I get seriously angry.

It hasn't happened in many months though so maybe they learned their lesson. I was so mad I seriously considered ending my Premium subscription right then and there when that track played.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#54
post #3

Thank you for sharing! On a tangent: I'm wondering if there are any good open source models/libraries to reconstruct audio quality. I'm thinking about an end-to-end open source alternative to something like Adobe Podcast [1] to make noisy recordings sound professional. Anecdotally it's supposed to be very good. In a recent search, I haven't found anything convincing. In my naive view this tasks seems much simpler tha…

It’s not open but Nvidia has RTX Voice for free if you have and Nvidia card.

Only weird thing it’s designed to be used real time but I’ve had some luck on cleaning up voice recordings replayed back through it via audio routing.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#55
post #29

It's interesting that the Death Metal was the hardest to reproduce. I conclude that it's the most fundamentally human of all genres.

The sound sample seemed to fit the 'vibe', but lacked any discernible definition. Could it be that it's too sonically dense to easily reproduce? Perhaps this could be improved with a more tailored training set.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#56
post #41
post #30

Earlier quoted context omitted.

I think the super weird part is that it's not great? I understand this is most likely very impressive technologically but musically it is disjointed, inconsistent and fake sounding. Most of the "music" examples have weird phrasing and confusing harmonic rhythm. Kudos to stability.ai for achieving this as I am sure it took a lot of effort and this is a huge leap forward in terms of generation of audio by generative AI…

I feel like music composition is a fundamentally hard task for AI. Music production seems like it should be a lot easier but I haven’t seen that

From what I've seen in the generated tracks so far (this one and others), they're pretty good locally, but just ignore the overall composition. For example any generated blues tracks will have the vague blues feel, but won't keep the 12 bar style. The bluegrass example here doesn't even seem to keep to 4/4 (or is extremely fluid about it...). Maybe one day someone will add a higher level "what's the current section, how far are you into it" inputs to that model to get something better - literally preparing the structure first and then filling it in. That should get much better results for context like "you're playing blues in A with quick change and generating bars 3-4, match the previous bars in style".

I mean, chatgpt knows how to plan this out https://chat.openai.com/share/976077c0-138b-4363-8065-3c8eed... Painting in that picture should be much easier than generating something freeflowing. Generating a good structure isn't that hard for most styles, because you can literally use the same pattern and do a few random changes that keep the key. (See lots of pop songs using the same 3/4 chord progression)

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#59
post #49
post #32

Earlier quoted context omitted.

Stability is great but Meta's MusicGen is available with code and weights while this isn't so that's a really odd place to make that comparison and complaint.

Before stable diffusion, nobody released weights at all. Meta et al only started sharing their models with the world when they realized how fast a developer ecosystem was building around the best models. Without stability, all of AI would still be closed and opaque.

>Before stable diffusion, nobody released weights at all.

That's not true. There's been a lot of models with weights from every player before Stability.

>Without stability, all of AI would still be closed and opaque.

Most GANs (the practically spiritual predecessor to diffusion models) for example were available. Huggingface existed and has realistically done more to keep AI open. And again, this specific release we are talking about by Stability is not Open.

Stability is great but you are re-writting history and doing it on the release where it makes least sense to do so.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#60
post #29

It's interesting that the Death Metal was the hardest to reproduce. I conclude that it's the most fundamentally human of all genres.

Well... they hardly tried all genres :)

It sounds like it can't handle lyrics or semantics that well so I suspect any genre where the lyricism is important would also be quite mushy and recognizably AI

Post reply on HN