Earlier quoted context omitted.
Wasn't this Fraunhofer's big insight that led to the development of MP3? Human perception actually is pretty forgiving of perturbations in the Fourier domain.
In very limited situations. You can move a frequency around (or drop it entirely) if it's being masked by a nearby loud frequency. Otherwise, you would be amazed at the sensitivity of pitch perception.
Riffusion – Stable Diffusion fine-tuned to generate music
371–380 of 481 posts
Re: Riffusion – Stable Diffusion fine-tuned to generate music
#372This really is unreasonably effective. Spectrograms are a lot less forgiving of minor errors than a painting. Move a brush stroke up or down a few pixels, you probably won't notice. Move a spectral element up or down a bit and you have a completely different sound. I don't understand how this can possibly be precise enough to generate anything close to a cohesive output. Absolutely blows my mind.
Author here: We were blown away too. This project started with a question in our minds about whether it was even possible for the stable diffusion model architecture to output something with the level of fidelity needed for the resulting audio to sound reasonable.
Re: Riffusion – Stable Diffusion fine-tuned to generate music
#373Re: Riffusion – Stable Diffusion fine-tuned to generate music
#374Quoted post unavailable.
Re: Riffusion – Stable Diffusion fine-tuned to generate music
#375Earlier quoted context omitted.
As a listener, I think you're probably still safe. Can you use this to help you though? Maybe. It's impressive what it produces, but I think it probably lacks substance in the same way the visual AI art stuff does. For the most part, it passes what I call the at-a-glanceness test. It's little better than apophenia (the same thing that makes you see shapes in clouds, faces in rocks, or think you've recognised a famili…
In general all this stuff is chopping the bottom off the market. AI art, code, writing, music, etc. can all generate passable "filler" content, which will decimate all human employment generating same. I don't think this stuff is a threat to genuinely innovative, thoughtful, meaningful work, but that's the top of the market. That being said the bottom of the market is how a lot of artists make their living, so this i…
> Basic income or revolution. That's going to be our choice.
I'm definitely pro basic income, but I've heard an interesting remark a few weeks ago. And that's that COVID was kind of a UBI experiment (in the US), albeit very limited, and it turned out that if people don't have to worry about making a living and don't have a job to work in then they'll start do stupid things on the internet. Like make up stupid conspiracy theories about vaccines. I can't remember who said this, it was one of the guests on Lex Fridman's podcast. I'm also not sure if it's a valid analogy but reminds me of Vonnegut's Player Piano.
Re: Riffusion – Stable Diffusion fine-tuned to generate music
#376Earlier quoted context omitted.
Author here: We were blown away too. This project started with a question in our minds about whether it was even possible for the stable diffusion model architecture to output something with the level of fidelity needed for the resulting audio to sound reasonable.
Any chance of spoken voice-work being possible? It would be interesting to see if a model could "speak" like James Earl Jones or Steve Blum.
Re: Riffusion – Stable Diffusion fine-tuned to generate music
#377I bet a cool riff on this would be to simply sample an ambient microphone in the workplace and use that the generate and slowly introduce matching background music that fits the current tenor of the environment. Done slowly and subtly enough I'd bet the listener may not even be entirely aware its happening. If we could measure certain kinds of productivity it might even be useful as a way to "extend" certain highly p…
Re: Riffusion – Stable Diffusion fine-tuned to generate music
#378prompt: relaxing jazz melody bass music negative_prompt: piano music
Re: Riffusion – Stable Diffusion fine-tuned to generate music
#379Earlier quoted context omitted.
Author here: We were blown away too. This project started with a question in our minds about whether it was even possible for the stable diffusion model architecture to output something with the level of fidelity needed for the resulting audio to sound reasonable.
Any chance of spoken voice-work being possible? It would be interesting to see if a model could "speak" like James Earl Jones or Steve Blum.
Steve Blum: https://fakeyou.com/tts/result/TR:xmjjq9ty0hnsyjrjnw806k6rnp...
Furiously working on voice-to-voice (web, real time, and singing!) Should be out the door tomorrow!