Live data from Hacker News

Riffusion – Stable Diffusion fine-tuned to generate music

riffusion.com

371–380 of 481 posts

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#371

Earlier quoted context omitted.

Wasn't this Fraunhofer's big insight that led to the development of MP3? Human perception actually is pretty forgiving of perturbations in the Fourier domain.

In very limited situations. You can move a frequency around (or drop it entirely) if it's being masked by a nearby loud frequency. Otherwise, you would be amazed at the sensitivity of pitch perception.

The easy example of this is playing a slightly out of tune guitar, or a mandolin where the strings in the course aren't matched in pitch perfectly. You can hear it, and it's just a few cents off.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#372
post #224

This really is unreasonably effective. Spectrograms are a lot less forgiving of minor errors than a painting. Move a brush stroke up or down a few pixels, you probably won't notice. Move a spectral element up or down a bit and you have a completely different sound. I don't understand how this can possibly be precise enough to generate anything close to a cohesive output. Absolutely blows my mind.

Author here: We were blown away too. This project started with a question in our minds about whether it was even possible for the stable diffusion model architecture to output something with the level of fidelity needed for the resulting audio to sound reasonable.

Any chance of spoken voice-work being possible? It would be interesting to see if a model could "speak" like James Earl Jones or Steve Blum.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#375
post #260

Earlier quoted context omitted.

As a listener, I think you're probably still safe. Can you use this to help you though? Maybe. It's impressive what it produces, but I think it probably lacks substance in the same way the visual AI art stuff does. For the most part, it passes what I call the at-a-glanceness test. It's little better than apophenia (the same thing that makes you see shapes in clouds, faces in rocks, or think you've recognised a famili…

In general all this stuff is chopping the bottom off the market. AI art, code, writing, music, etc. can all generate passable "filler" content, which will decimate all human employment generating same. I don't think this stuff is a threat to genuinely innovative, thoughtful, meaningful work, but that's the top of the market. That being said the bottom of the market is how a lot of artists make their living, so this i…

It gets better very quickly and we have no idea where its limitations are. In other words, we have no idea when the development will slow down significantly and how much of the bottom will it have chopped down by then. Whether it's 10% or maybe a 100.

> Basic income or revolution. That's going to be our choice.

I'm definitely pro basic income, but I've heard an interesting remark a few weeks ago. And that's that COVID was kind of a UBI experiment (in the US), albeit very limited, and it turned out that if people don't have to worry about making a living and don't have a job to work in then they'll start do stupid things on the internet. Like make up stupid conspiracy theories about vaccines. I can't remember who said this, it was one of the guests on Lex Fridman's podcast. I'm also not sure if it's a valid analogy but reminds me of Vonnegut's Player Piano.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#376
post #224

Earlier quoted context omitted.

Author here: We were blown away too. This project started with a question in our minds about whether it was even possible for the stable diffusion model architecture to output something with the level of fidelity needed for the resulting audio to sound reasonable.

Any chance of spoken voice-work being possible? It would be interesting to see if a model could "speak" like James Earl Jones or Steve Blum.

This already exists [1].

[1] https://www.respeecher.com/

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#377
post #44

I bet a cool riff on this would be to simply sample an ambient microphone in the workplace and use that the generate and slowly introduce matching background music that fits the current tenor of the environment. Done slowly and subtly enough I'd bet the listener may not even be entirely aware its happening. If we could measure certain kinds of productivity it might even be useful as a way to "extend" certain highly p…

Or perhaps use it in a hospital to play music that matches the state of a patient’s health as they are passing away.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#379
post #224

Earlier quoted context omitted.

Author here: We were blown away too. This project started with a question in our minds about whether it was even possible for the stable diffusion model architecture to output something with the level of fidelity needed for the resulting audio to sound reasonable.

Any chance of spoken voice-work being possible? It would be interesting to see if a model could "speak" like James Earl Jones or Steve Blum.

James Earl Jones: https://fakeyou.com/tts/result/TR:9ek4x6eb80kq49e94grnhctk4g...

Steve Blum: https://fakeyou.com/tts/result/TR:xmjjq9ty0hnsyjrjnw806k6rnp...

Furiously working on voice-to-voice (web, real time, and singing!) Should be out the door tomorrow!

Post reply on HN