Live data from Hacker News

Riffusion – Stable Diffusion fine-tuned to generate music

riffusion.com

31–40 of 481 posts

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#31

Earlier this year, graphic designers, last month it was software engineers, and now musicians are also feeling the effects. Who else will AI make looking for a new job?

The raw outputs of these tools will be best consumed by experts. Until general AI, these are just better tools for the same workers.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#32

This is a genius idea. Using an already-existing and well-performing image model, and just encoding input/output as a spectrogram... It's elegant, it's obvious in retrospective, it's just pure genius. I can't wait to hear some serious AI music-making a few years from now.

Makes me wonder if we will see a generalization of this idea. Just like in a CPU 90%+ of want you want to do can be modeled with very few instructions (mov, add, jmp..) we could see a set of very refined models (Stable difussion, GPT, etc) and all of their abstractions on top (ChatGPT, Rifussion, etc).

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#33

Earlier this year, graphic designers, last month it was software engineers, and now musicians are also feeling the effects. Who else will AI make looking for a new job?

Politicians, bureaucracy.

GPT-3, what policy should we apply to increase tax revenue by 5% given these constraints?

GPT-3, please tell me some populist thing to say to win the next election, or how should I deflect these corruption charges.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#34

This is a genius idea. Using an already-existing and well-performing image model, and just encoding input/output as a spectrogram... It's elegant, it's obvious in retrospective, it's just pure genius. I can't wait to hear some serious AI music-making a few years from now.

For what is worth, people were trying the same thing with GANs (I also played with doing it with stylegan a bit) but the results weren't as good.

The amazing thing is that the current diffusion models are so good that the spectograms are actually reasonable enough despite the small room for error.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#35

This really is unreasonably effective. Spectrograms are a lot less forgiving of minor errors than a painting. Move a brush stroke up or down a few pixels, you probably won't notice. Move a spectral element up or down a bit and you have a completely different sound. I don't understand how this can possibly be precise enough to generate anything close to a cohesive output. Absolutely blows my mind.

You can also add another neural-network to "smooth" the spectrogram, increase the resolution and remove artefacts, just like they do for image generation.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#36
post #3

Really cool. Can't get this to work on the homepage though. Might be a traffic thing? Edit: Works now. A bit laggy but it works. Brilliant!

I'm getting this back when I try to hear cats sing me a rock opera:

{"data":{"success":true,"worklet_output":{"error":"Model version 5qekv1q is not healthy"},"latency_ms":530}}

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#37
Some of this is really cool! The 20 step interpolations are very special, because they're concepts that are distinct and novel.

It absolutely sucks at cymbals, though. Everything sounds like realaudio :) composition's lacking, too. It's loop-y.

Set this up to make AI dubtechno or trip-hop. It likes bass and indistinctness and hypnotic repetitiveness. Might also be good at weird atonal stuff, because it doesn't inherently have any notion of what a key or mode is?

As a human musician and producer I'm super interested in the kinds of clarity and sonority we used to get out of classic albums (which the industry has kinda drifted away from for decades) so the way for this to take over for ME would involve a hell of a lot more resolution of the FFT imagery, especially in the highs, plus some way to also do another AI-ification of what different parts of the song exist (like a further layer but it controls abrupt switches of prompt)

It could probably do bad modern production fairly well even now :) exaggeration, but not much, when stuff is really overproduced it starts to get way more indistinct, and this can do indistinct. It's realaudio grade, it needs to be more like 128kbps mp3 grade.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#39

Earlier this year, graphic designers, last month it was software engineers, and now musicians are also feeling the effects. Who else will AI make looking for a new job?

Musicians were made to get a day job long before you were born ;)
Post reply on HN