This is a genius idea. Using an already-existing and well-performing image model, and just encoding input/output as a spectrogram... It's elegant, it's obvious in retrospective, it's just pure genius. I can't wait to hear some serious AI music-making a few years from now.
As someone who loves making music and loves listening to music made by other humans with intention, it just makes me sad. Sure, AI can do lots of things well. But would you rather live in a world where humans get to do things they love (and are able to afford a comfortable life while doing so) or a world where machines do the things humans love and humans are relegated to the remaining tasks that machines happened to…
Riffusion – Stable Diffusion fine-tuned to generate music
411–420 of 481 posts
Re: Riffusion – Stable Diffusion fine-tuned to generate music
#412Earlier quoted context omitted.
Considering Stable Diffusion generates 3-channel (RGB) images, maybe it would be possible to train it on amplitude and phase data as two different channels?
We took a look at encoding phase, but it is very chaotic and looks like Gaussian noise. The lack of spatial patterns is very hard for the model to generate. I think there are tons of promising avenues to improve quality though.
But the same thing could be done as a post-processing step, finding points where the spectrum is changing fast and resetting the phases to make a sharper transient.
Re: Riffusion – Stable Diffusion fine-tuned to generate music
#413Earlier quoted context omitted.
Considering Stable Diffusion generates 3-channel (RGB) images, maybe it would be possible to train it on amplitude and phase data as two different channels?
People have tried that, but the model essentially learns to discard the phase channel because it is too hard for it to learn any useful information from it.
Re: Riffusion – Stable Diffusion fine-tuned to generate music
#414Earlier quoted context omitted.
It is a Deep Truth in that the universe is predictable and can be represented (at least the parts we interact with) mathematically. Matrix algebra is a hellova a drug. I could imagine someone developing the ability to listen to spectrograms by looking at them.
There is a whole piece in Godel Escher Bach where they look at vinyl records as alll the soud data is in there.
Re: Riffusion – Stable Diffusion fine-tuned to generate music
#415Earlier quoted context omitted.
Any chance of spoken voice-work being possible? It would be interesting to see if a model could "speak" like James Earl Jones or Steve Blum.
James Earl Jones: https://fakeyou.com/tts/result/TR:9ek4x6eb80kq49e94grnhctk4g... Steve Blum: https://fakeyou.com/tts/result/TR:xmjjq9ty0hnsyjrjnw806k6rnp... Furiously working on voice-to-voice (web, real time, and singing!) Should be out the door tomorrow!
Re: Riffusion – Stable Diffusion fine-tuned to generate music
#416Earlier quoted context omitted.
Any chance of spoken voice-work being possible? It would be interesting to see if a model could "speak" like James Earl Jones or Steve Blum.
This already exists [1]. [1] https://www.respeecher.com/
I had a look around several months ago, and it seems like everything is locked behind SaaS APIs.
Re: Riffusion – Stable Diffusion fine-tuned to generate music
#417Earlier quoted context omitted.
I see a lot of AI naysayers neglecting the comparative advantage part. If AI completely eliminates low skill art labour from the job pool, it's not like those affected by it are gonna disintegrate, riot, and restructure society. They have the choice of filling an art niche an AI can't or they can spend that time learning other, more in-demand skills. This also ignores that fact that some companies would rather reallo…
I think specifically in the area of creative "products" such as art and music you have to think about the customer as well. I have zero interest in AI-created art or music. None. The value of art is its humanity; its expression of the artist's message, vision, and passion. AI doesn't have that, so it's not of any interest to me. I don't know how many custoners feel the same way, but I won't be purchasing any AI art o…
Re: Riffusion – Stable Diffusion fine-tuned to generate music
#418Also, spectrographs will never generate plausible high quality audio. (I think)
So I think the next move is to map the generate audio back over to synthesizer and samples via midi …
Re: Riffusion – Stable Diffusion fine-tuned to generate music
#419Re: Riffusion – Stable Diffusion fine-tuned to generate music
#420Earlier quoted context omitted.
> but the worst thing that can happen is that music will no longer be a profitable activity. For me, the worst that could happen is that people spend so much time listening to AI generated music, that human musicians can no longer find audiences to connect to. It's not just about economics (though that's also huge). It's the psychological cost of all of us spending greater and greater fractions of our lives connected…
Music was always about people. Even today, as most people listen to the mass-produced run-of-the-mill muzak, there is still a significant audience that seeks the "human element" for the sake of itself. Black metal community, for example, has always rejected all forms of "automation" and considers it not kvlt - rawness is a sought-after quality, defined as having people performing as close to the recording equipment a…