Live data from Hacker News

Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

stability.ai

91–100 of 210 posts

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#91
post #49
post #32

Earlier quoted context omitted.

Stability is great but Meta's MusicGen is available with code and weights while this isn't so that's a really odd place to make that comparison and complaint.

Before stable diffusion, nobody released weights at all. Meta et al only started sharing their models with the world when they realized how fast a developer ecosystem was building around the best models. Without stability, all of AI would still be closed and opaque.

In the image generation space, weights were never released for ImageGen and Dall-e, but yes you can find weights for more specialized generative models like StyleGAN (2, 3 etc). Stable Diffusion was arguably one of the most influential open model releases, and I think the substantial investment in StabilityAI is evidence of that.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#92
As an amateur musician, I’d be more interested in these tools if, along with the text description, they took as input a melody or chord progression or performance data. Maybe ABC notation or a MIDI track? Anyone doing that?

Other cool things would be a way to generate a sampled instrument from a text description, or to generate a new track given a text description and all the previous tracks for other instruments. There could be a new generation of audio tools that let you generate placeholders or better for everything.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#93
post #88

Earlier quoted context omitted.

Yessssss! I thought about MusicGAN and Markov chains last night thinking “Why can’t we just codify all chords and use a GAN to generate markov chains on chords of a key and have AI generate instruments and waveform from those chains?” IANA researcher but in my head, that sounded logical.

That's existed for decades. It's called Band in a Box. It's also cheezy as hell.

lol, no. Not autogenerate midi (although their latest versions of BiaB are pretty darn good now) but generate waveforms together. It would be similar to having AI generate whole scores of music but ensuring it's all in sync and in key. Not taking sample database of 88 sound files and triggering them when the midi-note strikes.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#94

Earlier quoted context omitted.

That raises an interesting difference between cleaning AI-generated sound and cleaning ordinary recordings. In an ordinary recording, there is an objective reality to discover -- a certain collection of voices was summed to create a signal. With (most? the best?) existing AI audio generation, the waveform is created from whole cloth, and extracting voices from it is an act of creation, not just discovery. I've come a…

Generative models can certainly create midi, but no one has done it yet. Given the technique is making video, audio, images, and language, all you need to do is train and build a model with an appropriate architecture. It’s easy to forget this is all pretty new stuff and it still costs a lot to make the base models. But the techniques are (more or less) well documented and implementable with open source tools.

> Generative models can certainly create midi, but no one has done it yet.

Note sequence generation from statistical models has a long history, at least as long if not longer than text generation.

Have a look at section 2.1 of this survey paper [0] that cites a paper from 1957 as the first work that applies Markov models to music generation.

And, of course, plenty of follow-up work 6 decades later on GANs, LSTMs, and transformers.

[0]: https://www.researchgate.net/publication/345915209_A_Compreh...

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#95
post #48

I would love to use it for background music when I am working. I have specific tastes that depend on the task, mood, energy level, and ambiance.

If you're not attuned to Cryo Chamber (label), check them out. Maybe not fitting all use-cases, but a strong and deep catalogue.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#96
post #32
post #4

I keep thinking back to when we didn't have stabilityai and it was just google and meta teasing us with mouth watering papers but never letting us touch them. I'm so thankful stability exists.

Stability is great but Meta's MusicGen is available with code and weights while this isn't so that's a really odd place to make that comparison and complaint.

Nah it's not because without the releases of Stability/ChatGPT it'd be the same situation. Cool nihilism though

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#97
post #59
post #49

Earlier quoted context omitted.

Before stable diffusion, nobody released weights at all. Meta et al only started sharing their models with the world when they realized how fast a developer ecosystem was building around the best models. Without stability, all of AI would still be closed and opaque.

>Before stable diffusion, nobody released weights at all. That's not true. There's been a lot of models with weights from every player before Stability. >Without stability, all of AI would still be closed and opaque. Most GANs (the practically spiritual predecessor to diffusion models) for example were available. Huggingface existed and has realistically done more to keep AI open. And again, this specific release we…

Nah. Dunno where this is coming from but infamously no AI models were released by big players for years. Rewind 18 months and all you got is GPT-3.0 that no one seems to care about and Disco Diffusion-y type stuff.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#98

As an amateur musician, I’d be more interested in these tools if, along with the text description, they took as input a melody or chord progression or performance data. Maybe ABC notation or a MIDI track? Anyone doing that? Other cool things would be a way to generate a sampled instrument from a text description, or to generate a new track given a text description and all the previous tracks for other instruments. Th…

The analogue from stable diffusion would be ControlNet, where you can train a superimposed model on auxiliary data, this should be possible to do with chords for example, just like you can do with human poses, 3D depth maps etc in stable diffusion using controlnet

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#99

Just like all examples of generative "AI" I've seen, there's always some bit of uncanny valley vibe present. In the audio examples, there's always this weird distortion like a really poorly compression sources were used as training data. The sounds are muddled together, and rarely do I hear clean musical voices. It's just a smear of sounds coming together that our brains try really hard to say "oh, that's a _____" si…

[deleted]

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#100

I still consider OpenAI's JukeBox (now at least 2 years old!) far and away the most creative music AI. But the combination of coherence, sound quality and creativity of this model is (to my knowledge) easily best in class.

The sound quality of Jukebox is muddled. There are many inconsistencies. The loudness of vocals and the quality of instruments really stand out and not in a good way. Hard to talk about creativity because it's so subjective but I've found it lacking in all AI music including JukeBox. Don't get me wrong - this tech is amazing.
Post reply on HN