Live data from Hacker News

Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

stability.ai

121–130 of 210 posts

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#121

Earlier quoted context omitted.

Midi is not lossy, midi is symbolic. There's a huge difference.

No, it's lossy. It's an event model at a fixed data rate. You can only do so many things sequentially, even if you could represent any possible musical concept as a MIDI event. So even if you're not sticking to note-on, note-off, it's still extremely lossy.

MIDI is able to accommodate nearly everything that can be represented through a musical score and instrumental performance. What are you hoping to accomplish with AI-generated waveforms that can't be done with MIDI?

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#122

It's interesting tech but none of the musical pieces impressed me (I play multiple instruments and have written and arranged music), most sounded too repetitive and not very imaginative. This is also an issue with diffusion based art AI in general, its good at a limited set of things but gets rather repetitive after a while. I could see using this as background music where quality is not important, like in games, tho…

>where quality is not important, like in games

I don't believe that's a good example. Video game music is an important part of the gaming experience, but its often taken for granted or overlooked.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#123

It's interesting tech but none of the musical pieces impressed me (I play multiple instruments and have written and arranged music), most sounded too repetitive and not very imaginative. This is also an issue with diffusion based art AI in general, its good at a limited set of things but gets rather repetitive after a while. I could see using this as background music where quality is not important, like in games, tho…

What I haven't seen done well by these generative AI's so far is structure (having a chorus, a verse, a bridge, ...) and harmonic movement/progressions (except for maybe a V-I or I - VI - ii - V) And those two be things are exactly what makes a song interesting and non-repetitive.

OpenAI's jukebox -- now 3 years old -- is creative and non-repetitive. Witness, for instance, its jam on Uptown Funk here:

https://www.youtube.com/watch?v=KCaya74_NHw

Or the changes shortly after 1:15, 2:15 and 2:40 in these extensions of Take On Me:

https://www.youtube.com/watch?v=_3yOrUJ0SzY

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#124
post #6

The bluegrass one is super weird. I can’t identify exactly why.

The AI seems to understand 4/4 time but doesn't understand groupings of 4 measures into phrases. It definitely doesn't understand ABABACA or even the basic parts of a song.

It is the musical equivalent of a meandering paragraph.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#126
post #49
post #32

Earlier quoted context omitted.

Stability is great but Meta's MusicGen is available with code and weights while this isn't so that's a really odd place to make that comparison and complaint.

Before stable diffusion, nobody released weights at all. Meta et al only started sharing their models with the world when they realized how fast a developer ecosystem was building around the best models. Without stability, all of AI would still be closed and opaque.

Stable Diffusion 1 contains a model OpenAI released. The CLIP encoder that was trained on text/image pairs at OpenAI.

https://huggingface.co/runwayml/stable-diffusion-v1-5

https://huggingface.co/runwayml/stable-diffusion-v1-5/blob/m...

Uploaded to Hugging Face Jan 2021

https://huggingface.co/openai/clip-vit-large-patch14

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#127
post #24

Earlier quoted context omitted.

Come on, be creative and make something new instead of copying someone else. It’s just kind of lame imo. “Mostly” private use? Mmm. :thumbs down emoji:

I meant private use and maybe share with a few friends. I actually agree with you that we probably shouldn't finetune on great artists and try to sell the output without modification or added creativity. Private or close friends sharing is fun and life enriching and inspiring though in my eyes.

Boards of Canada came to their sound because in their youth, the brothers had to move to Canada for a time. Even though it was only a couple of years their experience made an indelible mark on them—particularly school days watching old National Film Board of Canada tapes on worn VCR heads.

When they moved back to Scotland and started their music they started incorporating both the machinery and the sounds from the tapes in their compositions. And they could play their compositions live. It was quite the rig.

It’s not just entertainment. It’s communicating a very specific feeling and perspective. Keep learning and create, don’t be satisfied with just copying.

The biggest difference here is in the doing. You have to grow into one mode over time and energy spent, the other is immediate gratification with minimal personal energy.

Everything valuable comes during the course of that process of growing and committing energy. And it’s so good. Don’t deny yourself.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#128
post #59

Earlier quoted context omitted.

>Before stable diffusion, nobody released weights at all. That's not true. There's been a lot of models with weights from every player before Stability. >Without stability, all of AI would still be closed and opaque. Most GANs (the practically spiritual predecessor to diffusion models) for example were available. Huggingface existed and has realistically done more to keep AI open. And again, this specific release we…

Nah. Dunno where this is coming from but infamously no AI models were released by big players for years. Rewind 18 months and all you got is GPT-3.0 that no one seems to care about and Disco Diffusion-y type stuff.

You are looking at a very short part of recent history. It has not been like that at all.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#129
post #29

It's interesting that the Death Metal was the hardest to reproduce. I conclude that it's the most fundamentally human of all genres.

Well... they hardly tried all genres :) It sounds like it can't handle lyrics or semantics that well so I suspect any genre where the lyricism is important would also be quite mushy and recognizably AI

The Beatles seemed to be the hardest music for JukeBox to emulate.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#130
gamechanging stuff for sample based rap producers. havent been able to log-in yet but i think a good benchmark to start off with is to see if it can replicate the 'al green' sound from the early 70s - very distinct sounding production - drumless and instrumental.

you dont need 45 or 90 straight seconds of a coherent song rendered. just need to dip in the 45 sec clip and cut out 4 seconds here, another 4 there. reroll those cuts through stable audio, keep rolling, keep rolling. cut up and get a pile of clips together. arrange, layer, voila - you saved money on paying royalties for sampling.

the lofi melodic sample on the stability page was passable. thought the bluegrass one sounded great actually. imagine being able to program bluegrass like rap.

edit: oof. fully trained on a licensed commercial dataset from AudioSparx. muzak in, muzak out.

Post reply on HN