Live data from Hacker News

Riffusion – Stable Diffusion fine-tuned to generate music

riffusion.com

191–200 of 481 posts

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#191
post #22

They’ve got a looooong way to go man

I agree but it's better than listening to Ed Sheeran

Edit: To be honest, I find something like 'Band In A Box' to be more impressive and actually useful, I don't understand how I would ever use this or listen to this. To me, it's further proof that Stable Diffusion really just doesn't work that well

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#192
post #90

Some of this is really cool! The 20 step interpolations are very special, because they're concepts that are distinct and novel. It absolutely sucks at cymbals, though. Everything sounds like realaudio :) composition's lacking, too. It's loop-y. Set this up to make AI dubtechno or trip-hop. It likes bass and indistinctness and hypnotic repetitiveness. Might also be good at weird atonal stuff, because it doesn't inhere…

> composition's lacking, too. It's loop-y. Well no wonder, it has absolutely no concept of composition beyond a single 5s loop, if I understand correctly. > It absolutely sucks at cymbals, though. Everything sounds like realaudio :) > It could probably do bad modern production fairly well even now :) exaggeration, but not much, when stuff is really overproduced it starts to get way more indistinct, and this can do in…

No.

And I suspect this will always have phase smearing, because it's not doing any kind of source separation or individual synthesis. It's effectively a form of frequency domain data compression, so it's always going to be lossy.

It's more like a sophisticated timbral morph, done on a complete short loop instead of an individual line.

It would sound better with a much higher data density. CD quality would be 220500 samples for each five second loop. Realtime FFTs with that resolution aren't practical on the current generation of hardware, but they could be done in non-realtime. But there will always be the issue of timbres being distorted because outside of a certain level of familiarity and expectation our brains start hearing gargly disconnected overtones instead of coherent sound objects.

What this is not doing is extracting or understanding musical semantics and reassembling them in interesting ways. The harmonies in some of these clips are pretty weird and dissonant, and not what you'd get from a human writing accessible music. This matters because outside of TikTok music isn't about 5s loops, and longer structures aren't so amenable to this kind of approach.

This won't be a problem for some applications, but it's a long way short of the musical equivalent of a MidJourney image.

Generally we're a lot more tolerant of visual "bugs" than musical ones.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#193

Other author here! This got a posted a little earlier than we intended so we didn't have our GPUs scaled up yet. Please hang on and try throughout the day! Meanwhile, please read our about page http://riffusion.com/about It’s all open source and the code lives at https://github.com/hmartiro/riffusion-app --> if you have a GPU you can run it yourself This has been our hobby project for the past few months. Seeing the…

Amazing work. Can this be applied to voice?

Example prompt: “deep radio host voice saying ‘hello there’”

Kind of like a more expressive TTS?

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#195
post #38

This is really cool but can someone tell me why we are automating art? Who asked for this? The future seems depressing when I look at all this AI generated art.

I would say it's not "generated," but interpolated...

It doesn't make anything new or fresh. It doesn't pull any real-life emotions or experiences into a synthesis that a person can relate to. It's more like asking a teenaged comedian to imitate numerous impressions of music styles. e.g. in Clerks when the Russian guy does "metal": https://youtu.be/7gFoHkkCaRE?t=55

Of course the modern conception of music in the West is as an accompaniment to other, mostly drudging, activities, as opposed to something to be paid singularly attention. Therefore, there are many "valuable"(*) occasions to produce "impressions" of music. E.g. in advertisements and social media flexes where identity and attitude are the purpose of music. For these, a shallow interpretation or reflection of loosely amalgamated sound clips will suffice. But we don't just attend concerts or focus sustained energy on sonic impressions. We listen to lyrics and give over our consciousness to composed works because we want to find secrets others give away in dealing with this crazy thing called life- ideas to succeed, admissions of failure, and what the expected emotional arcs of these trajectories looks like. This lofty goal is to date not within the scope of AI stunts.

As Solzheinetysn said, "Too much art is like candy and not bread."

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#196
post #188
post #184

Earlier quoted context omitted.

You probably mean Karlheinz Brandenburg, the developer of MP3, who worked on psychoacoustics. Not completely off though, as he did the research at a Fraunhofer research institute, which takes its name from Joseph von Fraunhofer, the inventor of the spectroscope.

Does the institute not also claim that work?

Fair enough. But for me, when talking about `having an insight`, I don't imagine a non-human entity doing that. And to be pedantic (talking about Germans doing research, I hope everyone would expect me to be), the institute is called Fraunhofer IIS. `Fraunhofer` would colloquially refer to the society, which is an organization with 76 institutes total. Although, of course, the society will also claim the work...
Post reply on HN