Live data from Hacker News

Riffusion – Stable Diffusion fine-tuned to generate music

riffusion.com

381–390 of 481 posts

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#381
post #53

I find it really cool that the "uncanny valley" that's audible on nearly every sample is exactly as I would imagine that the visual artifacts would sound that crop up in most generated art. Not really surprising I guess, but still cool that there's such a direct correlation between completely different mediums!

yeah, it's pretty unsurprising, that they're both uncanny valley like a messy circus.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#382

Earlier quoted context omitted.

Any chance of spoken voice-work being possible? It would be interesting to see if a model could "speak" like James Earl Jones or Steve Blum.

James Earl Jones: https://fakeyou.com/tts/result/TR:9ek4x6eb80kq49e94grnhctk4g... Steve Blum: https://fakeyou.com/tts/result/TR:xmjjq9ty0hnsyjrjnw806k6rnp... Furiously working on voice-to-voice (web, real time, and singing!) Should be out the door tomorrow!

Excellent work! Singing would be amazing - karaoke can finally sound good :p

Have you released a tool for volumetric capture? I'm applying this to LED lighting fixture setup for tv/film/live shows and 3D positioning is the last step to fully automated configuration.

My goal is real-time sync between 3D model and real world.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#384

Earlier quoted context omitted.

Sadly, if we look at human history, it usually resolves to that.

But not always. And education and communication are some of the forces that can help avoid that. Knowledge is power.

education is controlled by government and communication is controlled by corporations

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#385

Earlier quoted context omitted.

I'm curious why, instead of using magnitude and phase, you wouldn't use real and imaginary parts?

There have been some attempts at doing this, some of which have been moderately successful. But fundamentally you still have the problem that from the NN's perspective, it's relatively easy for it to learn the magnitude but very hard for it to learn the phase. So it'll guess rough sizes for the real and imaginary parts, but it'll have a hard time learning the correct ratio between the two. Models which operate direct…

I wonder if there might be room for a hybrid approach, with a time-domain model taking machine-generated spectrograms as input and turning them into sound. (Just a thought, no idea whether it actually makes sense.)

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#386
post #208

This is a genius idea. Using an already-existing and well-performing image model, and just encoding input/output as a spectrogram... It's elegant, it's obvious in retrospective, it's just pure genius. I can't wait to hear some serious AI music-making a few years from now.

This idea is presented by Jeremy Howard on literally their first Deep Learning for Coders class (most recent edition). A student wanted to classify sounds, but only knew how to do vision, so they converted sounds to spectrograms, fine tuned the model on the labelled spectra, and the classification worked pretty well on test data. That of course does not take the merit away from the Riffusion authors though.

The idea to apply computer vision algorithms to spectrograms is not new. I don't know who first came up with it, but I first came across it about a decade ago.

I just ran a quick Google Scholar search, and the first result is https://ieeexplore.ieee.org/abstract/document/5672395

This is from 2010. I didn't go looking, but it wouldn't surprise me if the idea is older than that.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#387
post #377
post #44

I bet a cool riff on this would be to simply sample an ambient microphone in the workplace and use that the generate and slowly introduce matching background music that fits the current tenor of the environment. Done slowly and subtly enough I'd bet the listener may not even be entirely aware its happening. If we could measure certain kinds of productivity it might even be useful as a way to "extend" certain highly p…

Or perhaps use it in a hospital to play music that matches the state of a patient’s health as they are passing away.

I would not want to go to the hospital for a mild ear infection and hear the AI start blasting death metal.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#388

Earlier this year, graphic designers, last month it was software engineers, and now musicians are also feeling the effects. Who else will AI make looking for a new job?

Honestly none of them should. I think the moral panic around these things is way overstated. They are cool but hardly about to replace anyone's job.

Have you tried AI asset generators? They are working extremely good. Just yesterday a friend of mine has shown me the progress they made in their game. It is incredible. Designers are 100% loosing their job over this.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#389
post #247

Earlier quoted context omitted.

As one of the meatsacks whose job you're about to kill... eh, I got nothin, it's damn impressive. It's gonna hit electronic music like a nuclear bomb, I'd wager.

As a listener, I think you're probably still safe. Can you use this to help you though? Maybe. It's impressive what it produces, but I think it probably lacks substance in the same way the visual AI art stuff does. For the most part, it passes what I call the at-a-glanceness test. It's little better than apophenia (the same thing that makes you see shapes in clouds, faces in rocks, or think you've recognised a famili…

Potentially it could be used as temporary atmospherical music for pre-viz video shots

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#390

This really is unreasonably effective. Spectrograms are a lot less forgiving of minor errors than a painting. Move a brush stroke up or down a few pixels, you probably won't notice. Move a spectral element up or down a bit and you have a completely different sound. I don't understand how this can possibly be precise enough to generate anything close to a cohesive output. Absolutely blows my mind.

It's...not effective though. Am I listening to the wrong thing here? Everything I hear from the web app is jumbled nonsense.

Really? They sound quite clearly like the prompt to me if I “squint my ears” a little
Post reply on HN