Live data from Hacker News

Riffusion – Stable Diffusion fine-tuned to generate music

riffusion.com

431–440 of 481 posts

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#431

Earlier quoted context omitted.

Makes me wonder if we will see a generalization of this idea. Just like in a CPU 90%+ of want you want to do can be modeled with very few instructions (mov, add, jmp..) we could see a set of very refined models (Stable difussion, GPT, etc) and all of their abstractions on top (ChatGPT, Rifussion, etc).

Maybe next up is a model that generates Piet code https://www.dangermouse.net/esoteric/piet.html

and you ask stable diffusion to generate piet code for a slightly better version of stable diffusion (or chatGPT) ...which then you can further use to generate a better version, and so on. Singularity here we come!

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#432
post #106

Earlier quoted context omitted.

Wow, I am blown away. Some of these clips are really good! I love the Arabic Gospel one. John and George would have loved this so much. And the fact that you can make things that sound good by going through visual space feels to me like the discovery of a Deep Truth, one that goes beyond even the Fourier transform because it somehow connects the aesthetics of the two domains.

I can simultaneously burst a bubble and provide fuel for more -- the alignment of the intrinsic manifolds of different domains has been an interesting research topic for zero shot research for a few years. I remember seeing at CVPR 2018 the first zero shot...classifier, I think? That if I recall correctly trained in two domains that were automatically basically aligned with each other enough to provide very good zero…

Well it's no surprise that it kinda sorta works. Neural networks are very good at learning the underlying structure of things and working with suboptimally represented inputs. But if working with images of spectrograms works better than just samples in time domain, that is a valid and non-obvious finding.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#433

Other author here! This got a posted a little earlier than we intended so we didn't have our GPUs scaled up yet. Please hang on and try throughout the day! Meanwhile, please read our about page http://riffusion.com/about It’s all open source and the code lives at https://github.com/hmartiro/riffusion-app --> if you have a GPU you can run it yourself This has been our hobby project for the past few months. Seeing the…

What sort of setup do you need to be able to fine tune Stable Diffusion models? Are there good tutorials out there for fine tuning with cloud or non-cloud GPUs?

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#434

Earlier quoted context omitted.

"Horse Carriage Driver says horseless carriages are abominations. More at 12!"

adamsmith143, commenter. Witness here the apathy of someone unaffected by the suffering of millions and uncommitted to the soul of humanity. DIAF.

Sorry my sympathy for "starving" artists isn't sufficient for your liking. 95% of modern art is drivel that can and should be replaced by AI art. A recent walk through the MoMa in NY had me weeping for humanity. Filled with, in some cases literal, trash that any 5th grader could produce.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#435
post #247

Other author here! This got a posted a little earlier than we intended so we didn't have our GPUs scaled up yet. Please hang on and try throughout the day! Meanwhile, please read our about page http://riffusion.com/about It’s all open source and the code lives at https://github.com/hmartiro/riffusion-app --> if you have a GPU you can run it yourself This has been our hobby project for the past few months. Seeing the…

As one of the meatsacks whose job you're about to kill... eh, I got nothin, it's damn impressive. It's gonna hit electronic music like a nuclear bomb, I'd wager.

Isn't this just a sampler with extra steps?

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#436

Other author here! This got a posted a little earlier than we intended so we didn't have our GPUs scaled up yet. Please hang on and try throughout the day! Meanwhile, please read our about page http://riffusion.com/about It’s all open source and the code lives at https://github.com/hmartiro/riffusion-app --> if you have a GPU you can run it yourself This has been our hobby project for the past few months. Seeing the…

"fine-tuned on images of spectrograms paired with text"

How many paired training images / text and what was the source of your training data? Just curious to know how much fine tuning was needed to get the results and what the breadth / scope of the images were in terms of original sources to train on to get sufficient musical diversity.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#437
post #46

Authors here: Fun to wake up to this surprise! We are rushing to add GPUs so you can all experience the app in real-time. Will update asap

"fine-tuned on images of spectrograms paired with text"

How many paired training images / text and what was the source of your training data? Just curious to know how much fine tuning was needed to get the results and what the breadth / scope of the images were in terms of original sources to train on to get sufficient musical diversity.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#438
Is there a different mapping of FFT information to a two dimensional image that would make harmonic relationships more obvious?

IE, use a polar coordinate system where angle 12 oclock is 440hz, and the 12 chromatic notes would be mapped to the angle of the hours. Maybe red pixel intensity is bit mapped to octave, IE first, third and eight octave: 0b10100001.

Time would be represented by radius. Unfortunately the space wouldn't wrap nicely like if there was a native image format for representing donuts.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#439

This is a genius idea. Using an already-existing and well-performing image model, and just encoding input/output as a spectrogram... It's elegant, it's obvious in retrospective, it's just pure genius. I can't wait to hear some serious AI music-making a few years from now.

You already hear a ton of them. Lofi music on these massively popular channels are basically auto-generated "music" + auto generated artwork.

I dabble in music production and know some of the people in the "Lofi" world, so I know for a fact that this is not true. It's just a formulaic sub-genre where people are trying to make similar instrumentals with the same vibe. It would be jarring to listen to a playlist while studying and each song had wildly different tempos, instruments, etc.

Also, the music doesn't sound "Lofi" because it's generated by algorithms. A lot of hard work and software goes into taking a clean, pitch-perfect digital signal and making it sound like something playing on a record player from the 70s.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#440
post #220

the problem is it sounds awful, like a 64kbps MP3 or worse Perhaps AI can be trained to create music in different ways than generating spectrograms and converting them to audio?

It doesn't need to sound good at all for it to be useful. Like with the AI Art creation, it can be a starting point for artists to play around and rapidly try different concepts, and then interpret the concept using high quality tools to create something really quite remarkable. It's all about empowering artists to explore more possibilities.

I haven't been paying close attention to the AI field but seeing someone painting with NVIDIA Canvas blew me away[1].

I can totally see how it can help prototyping or exploring ideas and visions.

[1]: https://www.youtube.com/watch?v=uLlhyKxygrI

Post reply on HN