Live data from Hacker News

Stable-Audio-Demo

stability-ai.github.io

1–10 of 249 posts

Re: Stable-Audio-Demo

#3
So there aren't public weights, is that right? Having trouble finding anything that says one way or the other.

edit: Oh okay, didn't realize this was somehow a controversial comment to make. It would have been great if you had answered the question before downvoting but that's fine I suppose.

Re: Stable-Audio-Demo

#6
As with Stable Diffusion, text prompting will be the least controllable way to get useful output with this model. I can easily imagine midi being used as an input with control net to essentially get a neural synthesizer.

Re: Stable-Audio-Demo

#7
This is incredibly good compared to SOTA music models (MusicGen, MusicLM). It looks like there's also a product page where you can subscribe to use it, similar to Midjourney: https://www.stableaudio.com/

Sadly it's not open-weight and it doesn't look like there's an API (again like Midjourney): you subscribe monthly to generate audio in their UI, rather than having something developers can integrate or wrap.

Re: Stable-Audio-Demo

#8

As with Stable Diffusion, text prompting will be the least controllable way to get useful output with this model. I can easily imagine midi being used as an input with control net to essentially get a neural synthesizer.

Yes. Since working on my AI melodies project (https://www.melodies.ai/) two years ago, I've been saying that producing a high-quality, finalized song from text won't be feasible or even desirable for a while, and it's better to focus on using AI in various aspects of music making that support the artist's process.

Re: Stable-Audio-Demo

#9

This is right into the "uncanny valley" of music. It definitely sounded "like music", but none of it is what a human would produce. There's just something off.

The overall audio quality sounds pretty good and it seems to do a good job of sustaining a consistent rhythm and musical concept. But I agree there's something "off" about some of the clips.

- The rave music sounds great. But that's because EDM can be quite out there in terms of musical construction.

- The guitar sounds weird because it doesn't sound like chords a human hand can make on a tuning nobody tunes their guitar to - with a strange mix of open and closed strings that don't make sense. I think the restrictions of what a guitar can do aren't well understood by the model.

- The disco chord progression is bizarre. It doesn't sound bad, but it's unlikely to be something somebody working in the genre would choose.

- meditation music - I mean, most of that genre may as well just be some randomized process

- drum solo - there's some weird issues in some of the drum sounds, things like cymbals, rides and hats changing tone in the middle of a note, some of the toms sound weird, it sounds like a mix of stick and brush and stick and stick and brush all at the same time...it's sort of the same problem the solo guitar has where it's just not produced within the constraints of what a drum player can actually do on an instrument made of actual drums

- sound effects, all are pretty good, a little chunky and low bit-rate or low sample-rate sounding, there's probably something going on in the network that's reducing the rate before it gets build back up. There's a constant sort of reverb in all of the examples

I honestly can't say I prefer their model over some of the musicgen output even if their model is doing a better job at following the prompts in some cases.

All of the models have a very low bitrate encoding problems and other weird anomalous things. Some of it reminds me of the output from older mp3 encoders, where hihats and such would get very "swishy" sounding. You can hear some of it in the autoencoder reconstructions, especially the trumpet and the last example.

However, in any case, I'm actually glad in some ways to see the progress being made in this area. It's really impressive. This was complete science fiction only a very few years ago.

Post reply on HN