Live data from Hacker News

Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

stability.ai

181–190 of 210 posts

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#181

Earlier quoted context omitted.

No, it's lossy. It's an event model at a fixed data rate. You can only do so many things sequentially, even if you could represent any possible musical concept as a MIDI event. So even if you're not sticking to note-on, note-off, it's still extremely lossy.

MIDI is able to accommodate nearly everything that can be represented through a musical score and instrumental performance. What are you hoping to accomplish with AI-generated waveforms that can't be done with MIDI?

These diffusion models give an ability to create far more than can be communicated with MIDI + instrument. Riffusion gave a hint of this - rather than just notes and drum hits + some processing, it becomes one big pulsating, expressive mass which would not be reproducible without the granularity of a diffusion model. These are reminiscent of some of the serendipity of live recordings with lots of tracks where interesting things happen from interplay of many different layers. Generating a few dozen clips generally would give me 2 or 3 with a beautiful emotional passage which really lights up the pleasurable music part of my brain. Mass farming these clips seems like a good route to some amazing music.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#182

So.... Wait for llama for audio and train your own voice to having to call you friends by text and the software instead of actually saying the words? This is going to be nice for authentication, proving to a third party that you are yourself

Just wait for iOS17 https://youtu.be/oMt02DNbQlk

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#183
post #3

Thank you for sharing! On a tangent: I'm wondering if there are any good open source models/libraries to reconstruct audio quality. I'm thinking about an end-to-end open source alternative to something like Adobe Podcast [1] to make noisy recordings sound professional. Anecdotally it's supposed to be very good. In a recent search, I haven't found anything convincing. In my naive view this tasks seems much simpler tha…

https://youtu.be/o-kJ4_CuWzA

This video from MKBHD's studio channel dives into this topic

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#184

Earlier quoted context omitted.

Midi is not lossy, midi is symbolic. There's a huge difference.

MIDI is not inherently lossy. You could encode anything in it, just as you can encode any novel as an integer. In practice, though, transformations from audio to MIDI discard an enormous amount of important information, with the possible exception of transcriptions of performances on piano (where volume, frequency, duration and a good physical model of a piano are enough to reconstruct everything important about the…

MPE/MIDI 2.0 pretty much theoretically fixes the "discard an enormous amount of important information", in that you can have an extremely large number of parameters describing every single note and/or physical performance nuance.

Contemporary MIDI is ready to describe almost any human performance in an almost absurd level of detail. Whether anything can do something useful with the description is a different story.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#185

Earlier quoted context omitted.

Your example of the failures of MIDI are based on a 35-year old paper (from 1988!!!) about an earlier version of MIDI? When that paper was written it took several weeks and many millions of dollars of equipment to render primitive, mono-color 3d graphics. Desktop computers had 512 kilobytes of RAM and the highest-end desktops 32 MB of hard drive storage space. Computer screens had two colors: black and green. Audio c…

You clearly didn't read the article, and clearly don't understand how prevalent the MIDI 1.0 specification is today. MIDI 2.0 is a very recent development (this year LOL!) and has yet to be commercially adopted. The 1984 design is what is largely in use today. At the time of initial development, commercial synthesizers, not sound cards, were the intended generators of sound utilizing MIDI: https://www.vintagesynth.co…

MPE, which is the model for the most important part of MIDI 2.0, has been around for several years now; there are numerous software synths (and a few hardware ones) that can use it, and several hardware controllers ("instruments") that can deliver it.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#187

Earlier quoted context omitted.

What I haven't seen done well by these generative AI's so far is structure (having a chorus, a verse, a bridge, ...) and harmonic movement/progressions (except for maybe a V-I or I - VI - ii - V) And those two be things are exactly what makes a song interesting and non-repetitive.

OpenAI's jukebox -- now 3 years old -- is creative and non-repetitive. Witness, for instance, its jam on Uptown Funk here: https://www.youtube.com/watch?v=KCaya74_NHw Or the changes shortly after 1:15, 2:15 and 2:40 in these extensions of Take On Me: https://www.youtube.com/watch?v=_3yOrUJ0SzY

Okay but here it still seems like it's just taking over snippets. It did not come up with the progression.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#188

Earlier quoted context omitted.

> Can't name one art model you could run on your GPU from a FAANG or OpenAI before SD Google published dozens to promote Tensorflow: https://experiments.withgoogle.com/font-map https://experiments.withgoogle.com/sketch-rnn-demo https://experiments.withgoogle.com/curator-table https://experiments.withgoogle.com/nsynth-super https://experiments.withgoogle.com/t-sne-map The list goes on. Many are source-available with w…

I understand your point. The gap in communication is we don't mean _literally_ no one _ever_ open-sourced models. I agree, that would be absurd. [1] Companies, quite infamously and well-understood, _did_ hold back their "real" generative models, even from being available for pay. Take a stab at a literal definition: - post-GPT2 LLMs (ex. PALM, PALM2) - art like DaLL-E, Imagen, Parti Loosely, we had Disco Diffusion fo…

[deleted]

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#189
post #32
post #4

I keep thinking back to when we didn't have stabilityai and it was just google and meta teasing us with mouth watering papers but never letting us touch them. I'm so thankful stability exists.

Stability is great but Meta's MusicGen is available with code and weights while this isn't so that's a really odd place to make that comparison and complaint.

The way I see it regarding the point "but meta is also releasing models" is: there was one span of time between say 2014-2019 when mostly ML was just classifiers (nothing generative). People did open source those. Then there was a period between 2019-2023 when generative AI was possible. It's true that meta is releasing models in that space now finally. But there was an excruciating 3-4 year period between 2019 and 2022 when stable diffusion was finally made and released which opened the floodgates to others doing so as well. But I'm eternally grateful for emad and stabilitai for opening the gates that had been titillatingly closed for 4 annoying years.
Post reply on HN