Live data from Hacker News

Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

stability.ai

171–180 of 210 posts

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#171

Earlier quoted context omitted.

Literally nothing is preventing this other than that nobody has bothered to take the time to do it. The current version of MIDI is capable of replicating any of his performances, even down to the randomness. Note that if you want to replicate the audio quality of his performances, you will need a high-quality MIDI instrument; the ones that ship with Windows will not suffice. These MIDI instruments can range from a fe…

> nobody has bothered to take the time to do it In such case, we have a theoretical suggestion that «nothing is preventing this», but not an actual proof based on a "Turing test"-like scenario which would have specialists fooled, to corroborate that the new MIDI 2 would suffice.

I gave you the knowledge to do this research yourself, but since you are unwilling to do so, here is an example of the performance possible using MIDI instruments: https://www.youtube.com/watch?v=CvaChiq6gf0

As it is clear that you intend to keep shifting the goalposts to make a point that can't be made, I will withdraw from further participation in this thread.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#172

Earlier quoted context omitted.

Audio is definitely editable. While generative audio is new I am hopeful that a host of interesting applications will emerge (audio2audio etc.) within its ecosystem. Promising signal separation (audio to STEMs) and pitch detection tools already exist for raw audio signals. If you want to force Stability to focus on symbolic representations (such as severely lossy MIDI) I hope you can instead first try adapting to too…

Train the model with midi notes as text in the prompt and the audio as target. It will learn to interpret notes.

Not all music is well represented with notes, nor are audio datasets with high-quality note representations readily available. But I guess if you work hard enough you can get close: https://www.youtube.com/watch?v=o5aeuhad3OM My example still sounds like the chiptune simulation that it is, however.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#173

Earlier quoted context omitted.

The problems have long been known and articulated: http://www.music.mcgill.ca/~gary/courses/papers/Moore-Dysfun...

Your example of the failures of MIDI are based on a 35-year old paper (from 1988!!!) about an earlier version of MIDI? When that paper was written it took several weeks and many millions of dollars of equipment to render primitive, mono-color 3d graphics. Desktop computers had 512 kilobytes of RAM and the highest-end desktops 32 MB of hard drive storage space. Computer screens had two colors: black and green. Audio c…

You clearly didn't read the article, and clearly don't understand how prevalent the MIDI 1.0 specification is today. MIDI 2.0 is a very recent development (this year LOL!) and has yet to be commercially adopted. The 1984 design is what is largely in use today. At the time of initial development, commercial synthesizers, not sound cards, were the intended generators of sound utilizing MIDI: https://www.vintagesynth.com/roland/juno106.php.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#174

Earlier quoted context omitted.

No, it's lossy. It's an event model at a fixed data rate. You can only do so many things sequentially, even if you could represent any possible musical concept as a MIDI event. So even if you're not sticking to note-on, note-off, it's still extremely lossy.

MIDI is able to accommodate nearly everything that can be represented through a musical score and instrumental performance. What are you hoping to accomplish with AI-generated waveforms that can't be done with MIDI?

This is absurd. Sure, someone below posits that MIDI could perhaps represent a piano performance by Arturo Benedetti Michelangeli. I think it has been able to do a passing job at that, when you provide a decent piano. Regardless, piano rolls have been able to come close since the early 20th century. But how well does MIDI represent music performed by John Coltrane? Jimi Hendrix? It falls on its face. The long fetishized Western music notation abstraction, which MIDI poorly simulates, completely fails for many important examples of music. I would even venture to say that MIDI fails for most of them. But yes, MIDI is well-optimized for piano music where an acceptable piano or simulation is available.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#175

It's interesting tech but none of the musical pieces impressed me (I play multiple instruments and have written and arranged music), most sounded too repetitive and not very imaginative. This is also an issue with diffusion based art AI in general, its good at a limited set of things but gets rather repetitive after a while. I could see using this as background music where quality is not important, like in games, tho…

The real test is when this stuff is out in the wild and no one tells you it’s AI and the thought doesn’t cross your mind. Of course it’s not impressive nor surprising when the answer was given up front.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#176

Earlier quoted context omitted.

Generative models can certainly create midi, but no one has done it yet. Given the technique is making video, audio, images, and language, all you need to do is train and build a model with an appropriate architecture. It’s easy to forget this is all pretty new stuff and it still costs a lot to make the base models. But the techniques are (more or less) well documented and implementable with open source tools.

We’ve done it! wavtool.com

That’s really neat. How long have you been working on this?

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#177

It's interesting tech but none of the musical pieces impressed me (I play multiple instruments and have written and arranged music), most sounded too repetitive and not very imaginative. This is also an issue with diffusion based art AI in general, its good at a limited set of things but gets rather repetitive after a while. I could see using this as background music where quality is not important, like in games, tho…

[dead]

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#178

Earlier quoted context omitted.

We’ve done it! wavtool.com

That’s really neat. How long have you been working on this?

Thanks! It grew out of an old side project. Been full time on it since December.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#179
As a hobby musician, why don't they start with instrument samples? Every sampler user out there would love a button press generate sample on the fly as a plugin. It would blow away gigs and gigs of ridiculous duplicative or near duplicate samples.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#180

Earlier quoted context omitted.

Generative models can certainly create midi, but no one has done it yet. Given the technique is making video, audio, images, and language, all you need to do is train and build a model with an appropriate architecture. It’s easy to forget this is all pretty new stuff and it still costs a lot to make the base models. But the techniques are (more or less) well documented and implementable with open source tools.

> Generative models can certainly create midi, but no one has done it yet. Note sequence generation from statistical models has a long history, at least as long if not longer than text generation. Have a look at section 2.1 of this survey paper [0] that cites a paper from 1957 as the first work that applies Markov models to music generation. And, of course, plenty of follow-up work 6 decades later on GANs, LSTMs, and…

> cites a paper from 1957

By Fred Brooks no less…

https://en.m.wikipedia.org/wiki/Fred_Brooks

Post reply on HN