Live data from Hacker News

Riffusion – Stable Diffusion fine-tuned to generate music

riffusion.com

11–20 of 481 posts

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#13
Fun! I tried something similar with DCGAN when it first came out, but that didn't exactly make nice noises. The conversion to and from Mel spectrograms was lossy (to put it mildly), and DCGAN, while impressive in its day, is nothing like the stuff we have today.

Interesting that it gets so good results with just fine tuning the regular SD model. I assume most of the images it's trained on are useless for learning how to generate Mel spectrograms from text, so a model trained from scratch could potentially do even better.

There's still the issue of reconstructing sound from the spectrograms. I bet it's responsible for the somewhat tinny sound we get from this otherwise very cool demo.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#16
post #2

This is so good that I wondered if it's fake. Really impressive results from generated spectrographs! Also really interesting that it's not exactly trained on the audio files themselves - wonder if the usual copyright-based objections wild even apply here.

regarding those usual objections, i'd argue that a spectrograph representation of a given piece of audio is just a different (lossy) encoding of the same content/information, so any hypothetical objections would still apply here.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#18
This really is unreasonably effective. Spectrograms are a lot less forgiving of minor errors than a painting. Move a brush stroke up or down a few pixels, you probably won't notice. Move a spectral element up or down a bit and you have a completely different sound. I don't understand how this can possibly be precise enough to generate anything close to a cohesive output.

Absolutely blows my mind.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#19
This is a genius idea. Using an already-existing and well-performing image model, and just encoding input/output as a spectrogram... It's elegant, it's obvious in retrospective, it's just pure genius.

I can't wait to hear some serious AI music-making a few years from now.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#20
post #2

This is so good that I wondered if it's fake. Really impressive results from generated spectrographs! Also really interesting that it's not exactly trained on the audio files themselves - wonder if the usual copyright-based objections wild even apply here.

Why not? Music copyright was not even about audio recordings originally.
Post reply on HN