Live data from Hacker News

Riffusion – Stable Diffusion fine-tuned to generate music

riffusion.com

391–400 of 481 posts

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#391

Other author here! This got a posted a little earlier than we intended so we didn't have our GPUs scaled up yet. Please hang on and try throughout the day! Meanwhile, please read our about page http://riffusion.com/about It’s all open source and the code lives at https://github.com/hmartiro/riffusion-app --> if you have a GPU you can run it yourself This has been our hobby project for the past few months. Seeing the…

This is super awesome.

Have you already explored doing the same with voice cloning?

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#392

Other author here! This got a posted a little earlier than we intended so we didn't have our GPUs scaled up yet. Please hang on and try throughout the day! Meanwhile, please read our about page http://riffusion.com/about It’s all open source and the code lives at https://github.com/hmartiro/riffusion-app --> if you have a GPU you can run it yourself This has been our hobby project for the past few months. Seeing the…

funny that Hayk is an early skydio guy!

2 amazing AI projects. Huge respect :)

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#393
I feel like the next step here is to get a GPT3 like model to parse everything ever written about every piece of music which is on the internet (and in every pdf on libgen and scihub) and link them to spectrograms of that music

and then things are going to get wild

I am so blessed to live in this era :)

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#394
post #247

Earlier quoted context omitted.

As one of the meatsacks whose job you're about to kill... eh, I got nothin, it's damn impressive. It's gonna hit electronic music like a nuclear bomb, I'd wager.

As a listener, I think you're probably still safe. Can you use this to help you though? Maybe. It's impressive what it produces, but I think it probably lacks substance in the same way the visual AI art stuff does. For the most part, it passes what I call the at-a-glanceness test. It's little better than apophenia (the same thing that makes you see shapes in clouds, faces in rocks, or think you've recognised a famili…

I fully agree with what you wrote. This AI-generated music, while a great achievement, still sounds soulless. It's one thing to look at AI-generated pictures for a few seconds, but listening to this music with its gibberish "lyrics" for minutes really creeps me out - it's the "uncanny valley" all over again, I guess.

Regarding "can you use this to help you through?" - yeah, you could probably use it as a source of inspiration, but at the risk of getting sued by someone whos music you didn't even know you were copying...

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#395
post #66

This looks great and the idea is amazing. I tried with the prompt: "speed metal" and "speed metal with guitar riffs" and got some smooth rock-balad type music. I guess there was no heavy metal in the learning samples haha. Great work!

Gregorian death metal folk also seems to have lacked seed tunes but the thing is just in its infancy so soon we'll be banging our tonsured heads to the folky beats of ... ...OK, need to create a band name generator to work in tandem with this thing. Let's see what one of its brethren in ML makes of it... - "Echoes of the Past": This name plays on the idea of Gregorian chanting, which is often associated with the dist…

Somewhat unrelated, but given the descriptions, perhaps you should take a listen to the Darktide 40k soundtrack if you haven't: https://www.youtube.com/watch?v=D4hEOMSjzdo

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#396
post #224

Earlier quoted context omitted.

Author here: We were blown away too. This project started with a question in our minds about whether it was even possible for the stable diffusion model architecture to output something with the level of fidelity needed for the resulting audio to sound reasonable.

Any chance of spoken voice-work being possible? It would be interesting to see if a model could "speak" like James Earl Jones or Steve Blum.

have a look at UberDuck, they do something like this

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#397
post #66

This looks great and the idea is amazing. I tried with the prompt: "speed metal" and "speed metal with guitar riffs" and got some smooth rock-balad type music. I guess there was no heavy metal in the learning samples haha. Great work!

yeah, I also couldn't get it to do any folk metal. We shall have to wait for metal AI a short while yet haha

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#398

I'm a short fiction writer. Do you think I could get one of these new models to write a good story? I'd want to train it to include foreshadowing, suspense, relatable characters and perhaps a twist ending that cleverely references the beginning.

they can only do about 150 words at a time, so you'd struggle to get anything longform out of it. You can keep asking it over and over but then its liable to forget previous information. You'd do better with a prompt that repeatedly reminds the AI the style its going for and some basic information about the characters, but it's story would likely be quite cliche in many ways. It's something i've experimented with trying to get it to make DND adventure books

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#399

Earlier quoted context omitted.

Super! Makes sense since Skydio is also amazing. How much data is used for fine tuning? Since spectrograms are (surely?) very out of distribution for the pre training dataset, how much does value does the pre training really bring?

To be honest, we're not sure how much value image pre training brings. We have not tried to train from scratch, but it would be interesting. One thing that's very important though is the language pre-training. The model is able to do some amazing stuff with terms that do not appear in our data set at all. It does this by associating with related words that do appear in the dataset.

Hi, I really admire the skill you put at work on this project. At the same time, I think everyone is overlooking how crucial and problematic the training factor is.

Why was stable diffusion able to generate spectrograms? Because it was fed some. Presumably, those original spectrograms were scraped with little concern over creators' permissions, just like it has been for artists' work in order to produce art-looking image generation. Please, research what has been happening in the art community lately. https://www.youtube.com/watch?v=Nn_w3MnCyDY

A protest on ArtStation has been shown to influence Midjourney's results, proving that huge amounts of proprietary work are constantly scraped without the creators' permission. AIs like these work so well just because they steal and remix real artists' work in the first place. There are going to be legal wars about this.

Stable Diffusion doesn't have an official music generation Ai precisely because it couldn't train it with the same approach without being sued by music labels right away, while isolated artists don't have the same power.

So, back to my question: have you wondered whose work is Stable Diffusion remixing here? Your endeavour is great technically, but as we progress into the future we have to be more aware of the ethical implications that come with different forms of progress.

You could try to base your project on a collection of free-to-use spectograms, and see how it performs. If you do, I think it could actually be very interesting and useful to discuss the results here on Hacker News.

Cheers!

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#400

Earlier quoted context omitted.

As someone who loves making music and loves listening to music (regardless of its origins, in my case), it doesn't make me that sad. Sure, at first, I had an uncomfortable feeling that AI could make this sacred magic thing that only I and other fellow humans know how to do... But then I realized same thing is happening with visual art, so I applied the same counterarguments that've been cooking in my head. I think th…

> but the worst thing that can happen is that music will no longer be a profitable activity. For me, the worst that could happen is that people spend so much time listening to AI generated music, that human musicians can no longer find audiences to connect to. It's not just about economics (though that's also huge). It's the psychological cost of all of us spending greater and greater fractions of our lives connected…

Music was always about people. Even today, as most people listen to the mass-produced run-of-the-mill muzak, there is still a significant audience that seeks the "human element" for the sake of itself.

Black metal community, for example, has always rejected all forms of "automation" and considers it not kvlt - rawness is a sought-after quality, defined as having people performing as close to the recording equipment as possible.

There's also a rapper named Bones (Elmo O'Connor) who's never signed a contract with a label, does only music he wants to do, releases a couple albums every year. There's something about his approach that makes his music sound very organic and honest. I listen to him more than I listen to any mass produced rapper.

So in conclusion, music was always about people. Unless AI reaches AGI level, I don't think it will ever impact music enough to kill all audience.

Post reply on HN