Live data from Hacker News

Riffusion – Stable Diffusion fine-tuned to generate music

riffusion.com

1–10 of 481 posts

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#7
Producing images of spectrograms is a genius idea. Great implementation!

A couple of ideas that come to mind:

- I wonder if you could separate the audio tracks of each instrument, generate separately, and then combine them. This could give more control over the generation. Alignment might be tough, though.

- If you could at least separate vocals and instrumentals, you could train a separate model for vocals (LLM for text, then text to speech, maybe). The current implementation doesn't seem to handle vocals as well as TTS models.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#9
great stuff, while it comes with the usual smeary iFFT artifacts that AI-generated sound tends to have the results are surprisingly good. i especially love the nonsense vocals it generates in the last example, which remind me of what singing along to foreign songs felt like in my childhood.
Post reply on HN