I keep thinking back to when we didn't have stabilityai and it was just google and meta teasing us with mouth watering papers but never letting us touch them. I'm so thankful stability exists.
Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
101–110 of 210 posts
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#102Earlier quoted context omitted.
Having music editable for human post production is necessary for most professional adoption. Generating MIDIs would make much more sense than generating raw audio. This is what we do with AI images: you can fix them in Photoshop, etc. You cannot do this for raw audio due to how music is produced.
Build or seek out a MIDI generating model. I hope Stable Audio is never the place for that. MIDI is deeply lossy and it would be tragedy if it was the only music representation. Imagine if instead of phonographs, compact disks and streaming audio we only had piano rolls. What a loss indeed.
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#103Earlier quoted context omitted.
Generative models can certainly create midi, but no one has done it yet. Given the technique is making video, audio, images, and language, all you need to do is train and build a model with an appropriate architecture. It’s easy to forget this is all pretty new stuff and it still costs a lot to make the base models. But the techniques are (more or less) well documented and implementable with open source tools.
> Generative models can certainly create midi, but no one has done it yet. Note sequence generation from statistical models has a long history, at least as long if not longer than text generation. Have a look at section 2.1 of this survey paper [0] that cites a paper from 1957 as the first work that applies Markov models to music generation. And, of course, plenty of follow-up work 6 decades later on GANs, LSTMs, and…
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#104Earlier quoted context omitted.
Sir, your dog can talk! Yes, but not very well.
You joke, but even those videos of people saying their dog can talk is just like this. It's cute because it's a real dog making sounds we really want to believe when it's just them mimicking sounds because they get pettin's and treats. What I want is "AI" to do something impressive. Why are we trying to make the system generate the sounds itself? We don't make artists do that, we give them instruments. Give the model…
In the same way, many successful musicians can't read sheet music or know music theory, they just know how to produce something that sounds good.
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#105It's interesting tech but none of the musical pieces impressed me (I play multiple instruments and have written and arranged music), most sounded too repetitive and not very imaginative. This is also an issue with diffusion based art AI in general, its good at a limited set of things but gets rather repetitive after a while. I could see using this as background music where quality is not important, like in games, tho…
Yeah, the "rock drums" example was like a student in a practice. I'll be impressed when it can sound like Danny Carey. From all of the hype, I want to be impressed with results. Instead, we get these mediocre at best examples of what it can do. They are not good sales pitches to me.
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#106Thank you for sharing! On a tangent: I'm wondering if there are any good open source models/libraries to reconstruct audio quality. I'm thinking about an end-to-end open source alternative to something like Adobe Podcast [1] to make noisy recordings sound professional. Anecdotally it's supposed to be very good. In a recent search, I haven't found anything convincing. In my naive view this tasks seems much simpler tha…
We've been researching an audio denoiser for music that we will present at the AES conference in October. Description page: https://tape.it/denoising We'll also publish a webapp where you can use the denoiser for free. Mail me if you want beta access to it (email in profile). It won't be open-source though, although the paper will of course be public. It will also only reduce noise, and not reconstruct other aspects…
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#107Earlier quoted context omitted.
That's existed for decades. It's called Band in a Box. It's also cheezy as hell.
lol, no. Not autogenerate midi (although their latest versions of BiaB are pretty darn good now) but generate waveforms together. It would be similar to having AI generate whole scores of music but ensuring it's all in sync and in key. Not taking sample database of 88 sound files and triggering them when the midi-note strikes.
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#108Earlier quoted context omitted.
You joke, but even those videos of people saying their dog can talk is just like this. It's cute because it's a real dog making sounds we really want to believe when it's just them mimicking sounds because they get pettin's and treats. What I want is "AI" to do something impressive. Why are we trying to make the system generate the sounds itself? We don't make artists do that, we give them instruments. Give the model…
Well, we found with Midjourney et al that these models can work very well despite having no pre-conceived or symbolic notions of composition, color theory, perspective or anything. Yet they can produce really good results in the image generation space. It's the same idea here, except much earlier days. In the same way, many successful musicians can't read sheet music or know music theory, they just know how to produc…
Right, because they can operate the instruments that make the sound with natural talent, but they don't have to draw the waveforms. Audio generation is much different than image generation. It's just very odd to me.
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#109Earlier quoted context omitted.
Build or seek out a MIDI generating model. I hope Stable Audio is never the place for that. MIDI is deeply lossy and it would be tragedy if it was the only music representation. Imagine if instead of phonographs, compact disks and streaming audio we only had piano rolls. What a loss indeed.
Midi is not lossy, midi is symbolic. There's a huge difference.
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#110Earlier quoted context omitted.
Yeah, the "rock drums" example was like a student in a practice. I'll be impressed when it can sound like Danny Carey. From all of the hype, I want to be impressed with results. Instead, we get these mediocre at best examples of what it can do. They are not good sales pitches to me.
Hate to break it to you but there is a vast market for mediocre content, and even Danny Carey has a bad day once in a while.