Live data from Hacker News

Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

stability.ai

131–140 of 210 posts

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#131
post #35

Now imagine Spotify using this to generate individual earworms for everybody based on their personal tastes (likes, playlists). Yes, AI is partly hype, but had someone told me this even two years ago, I wouldn't have believed it.

Machine-generated music might be functionally equivalent to human-generated music, but that ignores the cultural role of art as a shared human experience - witness the liturgy of live music. That can't happen with music tailored to each listener, it can't happen without tracks that are fixed in time and can be referred to. I can imagine it well-accepted for dynamic music such as gaming soundtracks, but I suppose that…

Saying that's impossible makes me immediately wonder whether it's not. There are already headphone dance parties. What if a musical act's output was being interpreted through genre lenses specific to each listener?

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#132

I still consider OpenAI's JukeBox (now at least 2 years old!) far and away the most creative music AI. But the combination of coherence, sound quality and creativity of this model is (to my knowledge) easily best in class.

The sound quality of Jukebox is muddled. There are many inconsistencies. The loudness of vocals and the quality of instruments really stand out and not in a good way. Hard to talk about creativity because it's so subjective but I've found it lacking in all AI music including JukeBox. Don't get me wrong - this tech is amazing.

It's mushy and inconsistent, absolutely. But it also comes up with wild yet coherent changes that I've seen from nothing else.

At this comment I listed a few instances:

https://news.ycombinator.com/item?id=37499067

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#133

Earlier quoted context omitted.

Generative models can certainly create midi, but no one has done it yet. Given the technique is making video, audio, images, and language, all you need to do is train and build a model with an appropriate architecture. It’s easy to forget this is all pretty new stuff and it still costs a lot to make the base models. But the techniques are (more or less) well documented and implementable with open source tools.

> Generative models can certainly create midi, but no one has done it yet. Note sequence generation from statistical models has a long history, at least as long if not longer than text generation. Have a look at section 2.1 of this survey paper [0] that cites a paper from 1957 as the first work that applies Markov models to music generation. And, of course, plenty of follow-up work 6 decades later on GANs, LSTMs, and…

Do you know if anyone has tried training a text-to-music or text-to-midi model where the training data includes things like emotion labels for each note interval or chord progression?

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#134
post #38

Is the extreme metal music lacking from the training set? Why do the extreme metal examples always sound horrible?

Metal is especially hard to mix in a way that keeps the voices distinct and clear. Maybe the training catalog includes a lot of low-budget metal.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#135
post #6

The bluegrass one is super weird. I can’t identify exactly why.

The AI seems to understand 4/4 time but doesn't understand groupings of 4 measures into phrases. It definitely doesn't understand ABABACA or even the basic parts of a song. It is the musical equivalent of a meandering paragraph.

Absolutely. And all AI for music I have seen suffers that problem.

It makes me wonder whether the music generation should be stratified -- a coarse model lays out where parts like verse and chorus are, what distinguishes them, how to transition, etc., and then a finer-grained model fills in the details.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#136
post #41
post #30

Earlier quoted context omitted.

I think the super weird part is that it's not great? I understand this is most likely very impressive technologically but musically it is disjointed, inconsistent and fake sounding. Most of the "music" examples have weird phrasing and confusing harmonic rhythm. Kudos to stability.ai for achieving this as I am sure it took a lot of effort and this is a huge leap forward in terms of generation of audio by generative AI…

I feel like music composition is a fundamentally hard task for AI. Music production seems like it should be a lot easier but I haven’t seen that

Yes, and surprisingly so. I never would have guessed we'd have AI stock photographs before AI muzak.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#137

Earlier quoted context omitted.

> Generative models can certainly create midi, but no one has done it yet. Note sequence generation from statistical models has a long history, at least as long if not longer than text generation. Have a look at section 2.1 of this survey paper [0] that cites a paper from 1957 as the first work that applies Markov models to music generation. And, of course, plenty of follow-up work 6 decades later on GANs, LSTMs, and…

Do you know if anyone has tried training a text-to-music or text-to-midi model where the training data includes things like emotion labels for each note interval or chord progression?

That sounds expensive and inefficient. Peoples' interpretations of music (and abstract art more generally) can be shockingly different; I suspect the model would not get a clear signal from the result.

But that makes me wonder to what extent labeling can be programmed -- extracting chord changes, dynamics changes, tempo, gross timbral characteristics, etc.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#138

Earlier quoted context omitted.

Do you know if anyone has tried training a text-to-music or text-to-midi model where the training data includes things like emotion labels for each note interval or chord progression?

That sounds expensive and inefficient. Peoples' interpretations of music (and abstract art more generally) can be shockingly different; I suspect the model would not get a clear signal from the result. But that makes me wonder to what extent labeling can be programmed -- extracting chord changes, dynamics changes, tempo, gross timbral characteristics, etc.

And maybe even labels like popularity/play count/etc so it has a better sense of what “sounds good” to certain groups

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#139

Earlier quoted context omitted.

Nah. Dunno where this is coming from but infamously no AI models were released by big players for years. Rewind 18 months and all you got is GPT-3.0 that no one seems to care about and Disco Diffusion-y type stuff.

You are looking at a very short part of recent history. It has not been like that at all.

I'm all ears. I was "in the room" from 2019 on. Can't name one art model you could run on your GPU from a FAANG or OpenAI before SD, and can't name one LLM with public access before ChatGPT, much less weights available till LLaMA 1.

But please, do share.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#140
post #24

Earlier quoted context omitted.

I meant private use and maybe share with a few friends. I actually agree with you that we probably shouldn't finetune on great artists and try to sell the output without modification or added creativity. Private or close friends sharing is fun and life enriching and inspiring though in my eyes.

Boards of Canada came to their sound because in their youth, the brothers had to move to Canada for a time. Even though it was only a couple of years their experience made an indelible mark on them—particularly school days watching old National Film Board of Canada tapes on worn VCR heads. When they moved back to Scotland and started their music they started incorporating both the machinery and the sounds from the ta…

I get that. I think it's really cool what they did and when musicians put in time and energy into making amazing tracks. I get enough satisfaction from my normal coding job though, I don't have time to dedicate my life to music like they have. So from that perspective I'm just happy that it's possible to get more music like that type. Just a cool thing that exists in the world now, but I still think working hard to realize an artistic vision is also cool, separately.
Post reply on HN