Live data from Hacker News

Riffusion – Stable Diffusion fine-tuned to generate music

riffusion.com

221–230 of 481 posts

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#221

Earlier this year, graphic designers, last month it was software engineers, and now musicians are also feeling the effects. Who else will AI make looking for a new job?

This was the first AI thing to fill me with a feeling of existential dread.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#222
post #193

Other author here! This got a posted a little earlier than we intended so we didn't have our GPUs scaled up yet. Please hang on and try throughout the day! Meanwhile, please read our about page http://riffusion.com/about It’s all open source and the code lives at https://github.com/hmartiro/riffusion-app --> if you have a GPU you can run it yourself This has been our hobby project for the past few months. Seeing the…

Amazing work. Can this be applied to voice? Example prompt: “deep radio host voice saying ‘hello there’” Kind of like a more expressive TTS?

Author here: It can certainly be applied to voice, but the model would need deeper training to speak intelligibly. If you want to hear more singing, you can try a prompt like "female voice", and increase the denoising parameter in the settings of the app.

That said, our GPUs are still getting slammed today so you might face a delay in getting responses. Working on it!

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#223
post #220

the problem is it sounds awful, like a 64kbps MP3 or worse Perhaps AI can be trained to create music in different ways than generating spectrograms and converting them to audio?

It doesn't need to sound good at all for it to be useful. Like with the AI Art creation, it can be a starting point for artists to play around and rapidly try different concepts, and then interpret the concept using high quality tools to create something really quite remarkable.

It's all about empowering artists to explore more possibilities.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#224

This really is unreasonably effective. Spectrograms are a lot less forgiving of minor errors than a painting. Move a brush stroke up or down a few pixels, you probably won't notice. Move a spectral element up or down a bit and you have a completely different sound. I don't understand how this can possibly be precise enough to generate anything close to a cohesive output. Absolutely blows my mind.

Author here: We were blown away too. This project started with a question in our minds about whether it was even possible for the stable diffusion model architecture to output something with the level of fidelity needed for the resulting audio to sound reasonable.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#225
This is so completely wild. Love the novelty and inventiveness.

Could anyone help me understand whether using SVG instead of bitmap image would be possible? I realize that probably wouldn't be taking advantage of the current diffusion part of Stable-Diffusion, but my intuition is maybe it would be less noisy or offer a cleaner/more compressible path to parsing transitions in the latent space.

Great idea? Off base entirely? Would love some insight either way :D

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#226

I’m curious about the limitations of using spectrograms and transient-heavy sounds like drums. It seems like you’d need very high resolution spectrograms to get a consistently snappy drum sound.

8GB is enough to do 1080p resolution. the UI i use for SD maxes out at 2048x2048. however, it takes a lot longer than 512x512 to generate: 1m40s versus 1.97s. I'm guessing if one had access to one of those nvidia backplane rackmount devices one could generate 8k or larger resolution images.

SD can’t generate coherent images if you increase the output size. They’re basically always unusable unless you don’t need any global architecture to them.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#228
post #106

Other author here! This got a posted a little earlier than we intended so we didn't have our GPUs scaled up yet. Please hang on and try throughout the day! Meanwhile, please read our about page http://riffusion.com/about It’s all open source and the code lives at https://github.com/hmartiro/riffusion-app --> if you have a GPU you can run it yourself This has been our hobby project for the past few months. Seeing the…

Wow, I am blown away. Some of these clips are really good! I love the Arabic Gospel one. John and George would have loved this so much. And the fact that you can make things that sound good by going through visual space feels to me like the discovery of a Deep Truth, one that goes beyond even the Fourier transform because it somehow connects the aesthetics of the two domains.

I can simultaneously burst a bubble and provide fuel for more -- the alignment of the intrinsic manifolds of different domains has been an interesting research topic for zero shot research for a few years. I remember seeing at CVPR 2018 the first zero shot...classifier, I think? That if I recall correctly trained in two domains that were automatically basically aligned with each other enough to provide very good zero shot accuracy.

Calling it a Deep Truth might be a bit of an emotional marketing spin but the concept is very exciting nonetheless I believe.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#229
post #176

Pretty nice, I was just talking to a friend about needing a music version of chatgpt, so thank you for this. Wondering if it would be possible to create a version of this that you can point at a person SoundsCloud and have it emulate their style / create more music in the style of the original artist. I have a couple albums worth of downtempo electronic music I would love to point something like this at and see what…

https://mubert.com/ might be what you're looking for.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#230

Other author here! This got a posted a little earlier than we intended so we didn't have our GPUs scaled up yet. Please hang on and try throughout the day! Meanwhile, please read our about page http://riffusion.com/about It’s all open source and the code lives at https://github.com/hmartiro/riffusion-app --> if you have a GPU you can run it yourself This has been our hobby project for the past few months. Seeing the…

This is groundbreaking! All other attempts at AI generated music have IMO, fallen flat... These results are actually listenable, and enjoyable! This is almost frightening how powerful this can be
Post reply on HN