Live data from Hacker News

Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

stability.ai

81–90 of 210 posts

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#81

It's interesting tech but none of the musical pieces impressed me (I play multiple instruments and have written and arranged music), most sounded too repetitive and not very imaginative. This is also an issue with diffusion based art AI in general, its good at a limited set of things but gets rather repetitive after a while. I could see using this as background music where quality is not important, like in games, tho…

Yeah, the "rock drums" example was like a student in a practice. I'll be impressed when it can sound like Danny Carey.

From all of the hype, I want to be impressed with results. Instead, we get these mediocre at best examples of what it can do. They are not good sales pitches to me.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#82

It's interesting tech but none of the musical pieces impressed me (I play multiple instruments and have written and arranged music), most sounded too repetitive and not very imaginative. This is also an issue with diffusion based art AI in general, its good at a limited set of things but gets rather repetitive after a while. I could see using this as background music where quality is not important, like in games, tho…

I thought “Prompt: 116 BPM rock drums loop clean production” wasn't bad, but I'll grant that most of the rest would be an excellent way of showing (for instance) a death metal fan what their favoured music sounds like to those who don't have an ear for it :D

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#83

Earlier quoted context omitted.

That raises an interesting difference between cleaning AI-generated sound and cleaning ordinary recordings. In an ordinary recording, there is an objective reality to discover -- a certain collection of voices was summed to create a signal. With (most? the best?) existing AI audio generation, the waveform is created from whole cloth, and extracting voices from it is an act of creation, not just discovery. I've come a…

Generative models can certainly create midi, but no one has done it yet. Given the technique is making video, audio, images, and language, all you need to do is train and build a model with an appropriate architecture. It’s easy to forget this is all pretty new stuff and it still costs a lot to make the base models. But the techniques are (more or less) well documented and implementable with open source tools.

It has been done - first by OpenAI (MuseNet, which is no longer available) and later by Stanford (Anticipatory Music Transformer): https://nitter.net/jwthickstun/status/1669726326956371971

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#84

The solo piano was interesting because of how clean it is. I can imagine going from that sample to a score without too much difficulty. Once it's in a symbolic format it becomes much more flexible and re-usable. While this does not seem to be the trend I hope more gen ai in the audio and visual realms start to produce more structured / symbolic output. For example, if I were Adobe I would be training models, not to o…

Hopefully, the entire industry will NOT move in such a schematic and lossy direction. Use separate tools to analyze audio streams please. Don't throw the timbre baby out with the bathwater. MusicGen utilizes a tokenized transformer model for music, which is attractive for symbolic translation use cases. However, the overall audio quality is far more lossy than the examples you hear from Stable Audio. I believe that symbolic representation should not be a foundational approach to adequately represent and generate rich audio signals.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#85

It's interesting tech but none of the musical pieces impressed me (I play multiple instruments and have written and arranged music), most sounded too repetitive and not very imaginative. This is also an issue with diffusion based art AI in general, its good at a limited set of things but gets rather repetitive after a while. I could see using this as background music where quality is not important, like in games, tho…

Yeah, the "rock drums" example was like a student in a practice. I'll be impressed when it can sound like Danny Carey. From all of the hype, I want to be impressed with results. Instead, we get these mediocre at best examples of what it can do. They are not good sales pitches to me.

Sir, your dog can talk!

Yes, but not very well.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#86

The solo piano was interesting because of how clean it is. I can imagine going from that sample to a score without too much difficulty. Once it's in a symbolic format it becomes much more flexible and re-usable. While this does not seem to be the trend I hope more gen ai in the audio and visual realms start to produce more structured / symbolic output. For example, if I were Adobe I would be training models, not to o…

Having music editable for human post production is necessary for most professional adoption. Generating MIDIs would make much more sense than generating raw audio. This is what we do with AI images: you can fix them in Photoshop, etc. You cannot do this for raw audio due to how music is produced.

Build or seek out a MIDI generating model. I hope Stable Audio is never the place for that. MIDI is deeply lossy and it would be tragedy if it was the only music representation. Imagine if instead of phonographs, compact disks and streaming audio we only had piano rolls. What a loss indeed.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#87
post #32
post #4

I keep thinking back to when we didn't have stabilityai and it was just google and meta teasing us with mouth watering papers but never letting us touch them. I'm so thankful stability exists.

Stability is great but Meta's MusicGen is available with code and weights while this isn't so that's a really odd place to make that comparison and complaint.

Unfortunately MusicGen's output quality isn't strong enough. I applaud Meta for open sourcing it. The audio samples released for Stable Audio show much more promise. I look forward to code and model releases. I built out a Cog model for MusicGen and took it for a fairly extensive test drive and came back disappointed.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#88

The solo piano was interesting because of how clean it is. I can imagine going from that sample to a score without too much difficulty. Once it's in a symbolic format it becomes much more flexible and re-usable. While this does not seem to be the trend I hope more gen ai in the audio and visual realms start to produce more structured / symbolic output. For example, if I were Adobe I would be training models, not to o…

Yessssss! I thought about MusicGAN and Markov chains last night thinking “Why can’t we just codify all chords and use a GAN to generate markov chains on chords of a key and have AI generate instruments and waveform from those chains?” IANA researcher but in my head, that sounded logical.

That's existed for decades. It's called Band in a Box. It's also cheezy as hell.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#89

The solo piano was interesting because of how clean it is. I can imagine going from that sample to a score without too much difficulty. Once it's in a symbolic format it becomes much more flexible and re-usable. While this does not seem to be the trend I hope more gen ai in the audio and visual realms start to produce more structured / symbolic output. For example, if I were Adobe I would be training models, not to o…

I was wondering the same thing, definitely seems like generating the raw waveform runs into all kinds of weird issues (like they touched on in this post). I would imagine that training data would be a serious chokepoint here. Given how much discourse is currently kicking off around the intellectual property rights of just the final product (the mastered track), I can't imagine many musicians would be eager to share what is effectively the "proof of ownership" (track stems or MIDIs).

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#90

Earlier quoted context omitted.

Yeah, the "rock drums" example was like a student in a practice. I'll be impressed when it can sound like Danny Carey. From all of the hype, I want to be impressed with results. Instead, we get these mediocre at best examples of what it can do. They are not good sales pitches to me.

Sir, your dog can talk! Yes, but not very well.

You joke, but even those videos of people saying their dog can talk is just like this. It's cute because it's a real dog making sounds we really want to believe when it's just them mimicking sounds because they get pettin's and treats.

What I want is "AI" to do something impressive. Why are we trying to make the system generate the sounds itself? We don't make artists do that, we give them instruments. Give the models actual instruments, and then have it play them like a real artist. I will be much more impressed with an AI that understands composition and scoring, use of musical voices, key signatures. That would still be generative. I guess I just don't understand the point of the direction being taken. It's like a solution looking for a problem.

Post reply on HN