Live data from Hacker News

Riffusion – Stable Diffusion fine-tuned to generate music

riffusion.com

271–280 of 481 posts

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#271

This is a genius idea. Using an already-existing and well-performing image model, and just encoding input/output as a spectrogram... It's elegant, it's obvious in retrospective, it's just pure genius. I can't wait to hear some serious AI music-making a few years from now.

As someone who loves making music and loves listening to music made by other humans with intention, it just makes me sad. Sure, AI can do lots of things well. But would you rather live in a world where humans get to do things they love (and are able to afford a comfortable life while doing so) or a world where machines do the things humans love and humans are relegated to the remaining tasks that machines happened to…

As someone who loves making music and loves listening to music (regardless of its origins, in my case), it doesn't make me that sad. Sure, at first, I had an uncomfortable feeling that AI could make this sacred magic thing that only I and other fellow humans know how to do... But then I realized same thing is happening with visual art, so I applied the same counterarguments that've been cooking in my head.

I think that kind of attitude is defeatist - it's implying that humans will be stopped from making music if AI learns how to do it too. I don't think that will happen. Humans will continue making music, as they always have. When Kraftwerk started using computers to make music back in the 70s, people were also scared of what that will do to musicians. To be fair, live music has died out a bit (in a sense that there aren't that many god-on-earth-level rockstars), but it's still out there, people are performing, and others who want to listen can go and listen.

Maybe consumers will start consuming more and more AI music, instead of human music [0], but the worst thing that can happen is that music will no longer be a profitable activity. But then again, today's music industry already has some elements of the automation - washed-out rhythms, sexual thematics over and over again, re-hashing same old songs in different packages... So nothing's gonna change in the grand scheme of things.

[0] https://www.youtube.com/watch?v=S1jWdeRKvvk

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#272

Other author here! This got a posted a little earlier than we intended so we didn't have our GPUs scaled up yet. Please hang on and try throughout the day! Meanwhile, please read our about page http://riffusion.com/about It’s all open source and the code lives at https://github.com/hmartiro/riffusion-app --> if you have a GPU you can run it yourself This has been our hobby project for the past few months. Seeing the…

Hi Hayk, I see that the inference code and the final model are open source. I am not expecting it, but is the training code and the dataset you used for fine-tuning, and process to generate the dataset open source?

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#274
post #208

This is a genius idea. Using an already-existing and well-performing image model, and just encoding input/output as a spectrogram... It's elegant, it's obvious in retrospective, it's just pure genius. I can't wait to hear some serious AI music-making a few years from now.

This idea is presented by Jeremy Howard on literally their first Deep Learning for Coders class (most recent edition). A student wanted to classify sounds, but only knew how to do vision, so they converted sounds to spectrograms, fine tuned the model on the labelled spectra, and the classification worked pretty well on test data. That of course does not take the merit away from the Riffusion authors though.

The idea of connecting CV to audio via spectrograms pre dates Jeremy Howard's course by quite a bit. That's not really the interesting part here though. The fact that a simple extension of an image generation pipeline produces such impressive results with generative audio is what is interesting. It really emphasizes how useful the idea of stable diffusion is.

edit: added a bit more to the thought

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#275

Other author here! This got a posted a little earlier than we intended so we didn't have our GPUs scaled up yet. Please hang on and try throughout the day! Meanwhile, please read our about page http://riffusion.com/about It’s all open source and the code lives at https://github.com/hmartiro/riffusion-app --> if you have a GPU you can run it yourself This has been our hobby project for the past few months. Seeing the…

I've compiled/run a dozen different image to sound programs and none of them produce an acceptable sound. This bit of your code alone would be a great application by itself.

It'd be really cool if you could implement an MS paint style spectrum painting or image upload into the web app for more "manual" sound generation.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#276

Earlier quoted context omitted.

I can simultaneously burst a bubble and provide fuel for more -- the alignment of the intrinsic manifolds of different domains has been an interesting research topic for zero shot research for a few years. I remember seeing at CVPR 2018 the first zero shot...classifier, I think? That if I recall correctly trained in two domains that were automatically basically aligned with each other enough to provide very good zero…

It is a Deep Truth in that the universe is predictable and can be represented (at least the parts we interact with) mathematically. Matrix algebra is a hellova a drug. I could imagine someone developing the ability to listen to spectrograms by looking at them.

There is a whole piece in Godel Escher Bach where they look at vinyl records as alll the soud data is in there.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#277

Earlier quoted context omitted.

I've done some work on AI audio synthesis and the artifacts you're hearing in these clips are coming from the algorithm that is used to go from the synthesized spectrogram to the audio (the Griffin-Lim algorithm). Audio spectrograms have two components: the magnitude and the phase. Most of the information and structure is in the magnitude spectrogram so neural nets generally only synthesize that. If you were to look…

Phase is crtical for pitch. Here is why. The spectral transformation breaks up the signal into frequency bins. The frequency bins are not accurate enough to convey pitch properly. When a periodic signal is put through a FFT, it will land into a particular frequency bin. Say that the frequency of the signal is right in the middle of that bin. If you vary its pitch a little bit, it will still hand into the same bin. Kn…

Yes, this is exactly right and is why Griffin-Lim generated audio often has a sort of warbly quality. If you use a large FFT you can mitigate the issues with pitch because the frequency resolution in your spectrogram is higher, so the phase isn't so critical to getting the right pitch. But the trade-off of a bigger FFT is that the pitches now have to be stationary for longer.

The other place where phase is critical is in impulse sounds like drum beats. A short impulse is essentially just energy over a broad range of frequencies, but the phases have been chosen such that all the frequencies cancel each other out everywhere except for one short duration where they all add constructively. Without the right phases, these kinds of sounds get smeared out in time and sound sort of flat and muffled. The typing example on their demo page is actually a good example of this.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#278
Awesome work.

Would you be willing to share details about the fine-tuning procedure, such as the initialization, learning rate schedule, batch size, etc.? I'd love to learn more.

Background: I've been playing around with generating image sequences from sliding windows of audio. The idea roughly works, but the model training gets stuck due to the difficulty of the task.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#280
post #260

Earlier quoted context omitted.

As a listener, I think you're probably still safe. Can you use this to help you though? Maybe. It's impressive what it produces, but I think it probably lacks substance in the same way the visual AI art stuff does. For the most part, it passes what I call the at-a-glanceness test. It's little better than apophenia (the same thing that makes you see shapes in clouds, faces in rocks, or think you've recognised a famili…

In general all this stuff is chopping the bottom off the market. AI art, code, writing, music, etc. can all generate passable "filler" content, which will decimate all human employment generating same. I don't think this stuff is a threat to genuinely innovative, thoughtful, meaningful work, but that's the top of the market. That being said the bottom of the market is how a lot of artists make their living, so this i…

The only thing that affects whether you have a job is the Federal Reserve, not how good productivity tools are. You always have comparative advantage vs an AI, so you always have the qualifications for an entry level job.

There will never be a revolution and there's no such thing as late capitalism. Well, not if the Fed does their job.

Post reply on HN