Earlier quoted context omitted.
Stability is great but Meta's MusicGen is available with code and weights while this isn't so that's a really odd place to make that comparison and complaint.
Before stable diffusion, nobody released weights at all. Meta et al only started sharing their models with the world when they realized how fast a developer ecosystem was building around the best models. Without stability, all of AI would still be closed and opaque.
Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
91–100 of 210 posts
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#92Other cool things would be a way to generate a sampled instrument from a text description, or to generate a new track given a text description and all the previous tracks for other instruments. There could be a new generation of audio tools that let you generate placeholders or better for everything.
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#93Earlier quoted context omitted.
Yessssss! I thought about MusicGAN and Markov chains last night thinking “Why can’t we just codify all chords and use a GAN to generate markov chains on chords of a key and have AI generate instruments and waveform from those chains?” IANA researcher but in my head, that sounded logical.
That's existed for decades. It's called Band in a Box. It's also cheezy as hell.
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#94Earlier quoted context omitted.
That raises an interesting difference between cleaning AI-generated sound and cleaning ordinary recordings. In an ordinary recording, there is an objective reality to discover -- a certain collection of voices was summed to create a signal. With (most? the best?) existing AI audio generation, the waveform is created from whole cloth, and extracting voices from it is an act of creation, not just discovery. I've come a…
Generative models can certainly create midi, but no one has done it yet. Given the technique is making video, audio, images, and language, all you need to do is train and build a model with an appropriate architecture. It’s easy to forget this is all pretty new stuff and it still costs a lot to make the base models. But the techniques are (more or less) well documented and implementable with open source tools.
Note sequence generation from statistical models has a long history, at least as long if not longer than text generation.
Have a look at section 2.1 of this survey paper [0] that cites a paper from 1957 as the first work that applies Markov models to music generation.
And, of course, plenty of follow-up work 6 decades later on GANs, LSTMs, and transformers.
[0]: https://www.researchgate.net/publication/345915209_A_Compreh...
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#95I would love to use it for background music when I am working. I have specific tastes that depend on the task, mood, energy level, and ambiance.
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#96I keep thinking back to when we didn't have stabilityai and it was just google and meta teasing us with mouth watering papers but never letting us touch them. I'm so thankful stability exists.
Stability is great but Meta's MusicGen is available with code and weights while this isn't so that's a really odd place to make that comparison and complaint.
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#97Earlier quoted context omitted.
Before stable diffusion, nobody released weights at all. Meta et al only started sharing their models with the world when they realized how fast a developer ecosystem was building around the best models. Without stability, all of AI would still be closed and opaque.
>Before stable diffusion, nobody released weights at all. That's not true. There's been a lot of models with weights from every player before Stability. >Without stability, all of AI would still be closed and opaque. Most GANs (the practically spiritual predecessor to diffusion models) for example were available. Huggingface existed and has realistically done more to keep AI open. And again, this specific release we…
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#98As an amateur musician, I’d be more interested in these tools if, along with the text description, they took as input a melody or chord progression or performance data. Maybe ABC notation or a MIDI track? Anyone doing that? Other cool things would be a way to generate a sampled instrument from a text description, or to generate a new track given a text description and all the previous tracks for other instruments. Th…
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#99Just like all examples of generative "AI" I've seen, there's always some bit of uncanny valley vibe present. In the audio examples, there's always this weird distortion like a really poorly compression sources were used as training data. The sounds are muddled together, and rarely do I hear clean musical voices. It's just a smear of sounds coming together that our brains try really hard to say "oh, that's a _____" si…
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#100I still consider OpenAI's JukeBox (now at least 2 years old!) far and away the most creative music AI. But the combination of coherence, sound quality and creativity of this model is (to my knowledge) easily best in class.