Live data from Hacker News

Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

stability.ai

111–120 of 210 posts

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#111
Humans take a long time to get good at art; in the meantime they still have to eat.

So they compete with generative AI for a fixed number of jobs. The AI is cheaper and faster. Humans stop training to become artists.

Without new training data, the generative AI models stagnate. Progress in art stops globally, forever.

But for a brief glorious moment, we were able to say "huh, that's not bad".

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#112
post #98

As an amateur musician, I’d be more interested in these tools if, along with the text description, they took as input a melody or chord progression or performance data. Maybe ABC notation or a MIDI track? Anyone doing that? Other cool things would be a way to generate a sampled instrument from a text description, or to generate a new track given a text description and all the previous tracks for other instruments. Th…

The analogue from stable diffusion would be ControlNet, where you can train a superimposed model on auxiliary data, this should be possible to do with chords for example, just like you can do with human poses, 3D depth maps etc in stable diffusion using controlnet

It’s coming

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#114
I wonder if it makes sense to generate a combo of instruments rather than individual voices and then combine those with an arranger DNN. I would think it'd be much easier to capture each instrument's transients and dynamics that way, much less allow more subtlety in how they combine, like allowing the lead voice to shift among instruments, or even let the listener choose how each voice expresses stylistically and how they should combine.

Trying to do all of that in a single DNN, much less parameterize it useably seems overly ambitious (or will be of more limited value ultimately).

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#115

Humans take a long time to get good at art; in the meantime they still have to eat. So they compete with generative AI for a fixed number of jobs. The AI is cheaper and faster. Humans stop training to become artists. Without new training data, the generative AI models stagnate. Progress in art stops globally, forever. But for a brief glorious moment, we were able to say "huh, that's not bad".

This is by design - the capital that backs modern art isn’t doing it for love of the art but for money.

For fine art, it’s a way for them to launder money and keep it out of bank accounts where it can be seized trivially.

For mass art, it’s about selling to enough rubes to make a profit.

Neither are impacted by a stagnation in art. If anything, they’re aided by it - suddenly the art you bought to launder money retains its value because it’s no longer the flavor of the week with the arts crowd.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#117

Earlier quoted context omitted.

Build or seek out a MIDI generating model. I hope Stable Audio is never the place for that. MIDI is deeply lossy and it would be tragedy if it was the only music representation. Imagine if instead of phonographs, compact disks and streaming audio we only had piano rolls. What a loss indeed.

Midi is not lossy, midi is symbolic. There's a huge difference.

MIDI is not inherently lossy. You could encode anything in it, just as you can encode any novel as an integer.

In practice, though, transformations from audio to MIDI discard an enormous amount of important information, with the possible exception of transcriptions of performances on piano (where volume, frequency, duration and a good physical model of a piano are enough to reconstruct everything important about the signal) and similar instruments.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#118

Does this model support / "understand" concepts of spatial audio? For example, something like "an alarm moving around you in a circle". When AudioGen was announced this was my first question, but from what I've been able to test the model just ignores spatial audio prompts. Unfortunately I haven't been able to find any discussion or interest in online discussion about the importance / significance of spatial audio. W…

My guess is that it's not a very interesting problem because it's not particularly difficult to add spatial dimensions to arbitrary audio - after all, it is already commonly done in video games. All you have to do is manipulate the multichannel outputs with an understanding of the spatial positioning of each channel's speaker location relative to the listener and some basic trig.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#119

It's interesting tech but none of the musical pieces impressed me (I play multiple instruments and have written and arranged music), most sounded too repetitive and not very imaginative. This is also an issue with diffusion based art AI in general, its good at a limited set of things but gets rather repetitive after a while. I could see using this as background music where quality is not important, like in games, tho…

I disagree with your diffusion based art assessment, and I think it’s probably colored by what most people seem to want to make with it. Just as with regular art, you need some sort of vision to go beyond what everyone else does. Prompting “pretty girl wearing sexy clothing” for the umpteenth time isn’t new.

AI art gets rid of the technical skill step but the rest is there, although you may luck in to something at random. If you’re using ControlNet on Stable Diffusion or training your own models you have a lot of control over the output as well.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#120

Earlier quoted context omitted.

Sir, your dog can talk! Yes, but not very well.

You joke, but even those videos of people saying their dog can talk is just like this. It's cute because it's a real dog making sounds we really want to believe when it's just them mimicking sounds because they get pettin's and treats. What I want is "AI" to do something impressive. Why are we trying to make the system generate the sounds itself? We don't make artists do that, we give them instruments. Give the model…

We work with what we have. We don't have a lot of recordings of the physical movements of musicians; we have recordings.

Similarly, we don't have recordings of the actions of painters; we have finished paintings -- but if you're not impressed with what AI can do in the visual sphere, your standards are, to put it mildly, high.

Post reply on HN