Live data from Hacker News

Riffusion – Stable Diffusion fine-tuned to generate music

riffusion.com

351–360 of 481 posts

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#351
post #93

Other author here! This got a posted a little earlier than we intended so we didn't have our GPUs scaled up yet. Please hang on and try throughout the day! Meanwhile, please read our about page http://riffusion.com/about It’s all open source and the code lives at https://github.com/hmartiro/riffusion-app --> if you have a GPU you can run it yourself This has been our hobby project for the past few months. Seeing the…

All the AI music I’ve heard so far has a really unpleasant resonant quality to it. Why is that? Can it be removed?

The first ever recordings had people shouting to get anything to register. They sounded like tin. Fast forward to today.

Looking back at image generation just a year or two ago and people would have said similar things.

Not hard to imagine the trajectory of synthesized audio taking a similar path.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#352
I wonder if subliminal messaging will somehow make a comeback once we have ai generated audio and video. Something like we type "Fun topic" and those controlling the servers will inject "and praise to our empire/government/tyrant" to the suggestion or something like that.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#353
This opens up ideas. One thing people have tried to do with stable diffusion is create animations. Of course, they all come out pretty janky and gross, you can't get the animation smooth.

But what if what if a model was trained not on single images, but animated sequential frames, in sets, laid out on a single visual plane. So a panel might show a short sequence of a disney princess expressing a particular emotion as 16 individual frames collected as a single image. One might then be able to generate a clean animated sequence of a previously unimagined disney princess expressing any emotion the model has been trained on. Of course, with big enough models one could (if they can get it working) produced text prompted animations across a wide variety of subjects and styles.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#354

This really is unreasonably effective. Spectrograms are a lot less forgiving of minor errors than a painting. Move a brush stroke up or down a few pixels, you probably won't notice. Move a spectral element up or down a bit and you have a completely different sound. I don't understand how this can possibly be precise enough to generate anything close to a cohesive output. Absolutely blows my mind.

It's...not effective though. Am I listening to the wrong thing here? Everything I hear from the web app is jumbled nonsense.

Look at GAN art from a few years ago, compared to MidJourney v4.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#356

Earlier quoted context omitted.

I see a lot of AI naysayers neglecting the comparative advantage part. If AI completely eliminates low skill art labour from the job pool, it's not like those affected by it are gonna disintegrate, riot, and restructure society. They have the choice of filling an art niche an AI can't or they can spend that time learning other, more in-demand skills. This also ignores that fact that some companies would rather reallo…

I think specifically in the area of creative "products" such as art and music you have to think about the customer as well. I have zero interest in AI-created art or music. None. The value of art is its humanity; its expression of the artist's message, vision, and passion. AI doesn't have that, so it's not of any interest to me. I don't know how many custoners feel the same way, but I won't be purchasing any AI art o…

I think it's an interesting perspective but I will be very surprised if it's one that is common when this becomes more of a real choice. If there's two mp3s and one of them is more enjoyable to listen to, very few people will stick with the song they enjoy less because it's not AI generated.

Maybe a parallel would be furniture; there are people who buy hand crafted furniture but it's kind of a luxury. Most people just have Ikea and wouldn't pay more for the same (or have less good furniture) just to get some artisanal dinner table chair.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#357

I'm a short fiction writer. Do you think I could get one of these new models to write a good story? I'd want to train it to include foreshadowing, suspense, relatable characters and perhaps a twist ending that cleverely references the beginning.

I’ve thought about this and tested it a lot, it doesn’t work.

The reason is, the models don’t understand content or even context, they just recognize patterns and can generate similar patterns which we then interpret.

Case in point, a photo of an Astronaut on a horse is not actually an astronaut on a horse, it’s a 2D pixel map of light that our eyes then interpret to mean an astronaut on a horse. It’s even easier to understand this listening to the generated singing. It sounds just like singing - but it’s not language at all and doesn’t mean anything, it’s just a very similar pattern.

These AI models are great at making things that look like a pattern that to us means something, but they don’t generate semantically meaningful content itself.

So when you do this with GPT3 or other models and long-form narrative it falls apart pretty fast as the model can’t keep straight things like characters and their internal personalities and motives, nor the overall plot arc.

But! You can definitely feed in prompts and ask questions to get ideas and boilerplate - descriptions of people and places come out especially well - and then you can edit that and use it to accelerate your writing process.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#359
post #247

Earlier quoted context omitted.

As one of the meatsacks whose job you're about to kill... eh, I got nothin, it's damn impressive. It's gonna hit electronic music like a nuclear bomb, I'd wager.

These are tools. Don't think of them as replacements, they aren't. But as tools that will help us be creative. As smart as these apps seem, they will still need a human to decide where and how to use them. They won't replace us but we need to adapt to a new reality.

I hear this a lot (in relation to various jobs) and I still don't get it. Yes, it is a tool. Yes, if it can, it will replace humans. That's the whole point.

For some reason people tend to think that these tools/AI/ML systems will never be good enough to do their job (or a specific job). This argument can take different forms, sometimes stating that it will just do the boring part of the work (e.g. with programming) or that it will still need human creativity (maybe, but not necessarily and that's not the point) or that it will just replace low level, unskilled or mediocre professionals. And somehow everyone thinks they are not mediocre (i.e. average). But even these assumptions are unfounded. Why would anyone think that these systems will top out below their skill levels? Why would anyone think that they can't become superhuman?

They did in chess, go, I think poker too. Not to mention protein folding. And without much of a hitch between mediocre/good enough and superhuman. Because that difference is just interesting for us, but doesn't necessarily mean that there is huge step, that the system needs to undergo serious development and that it would take a long time. (Like decades or so.) People thought that was the case when AlphaGo beat Fan Hui saying that Lee Sedol was a completely different level. Which, of course, he is. Still, it just took DeepMind half a year to improve alphago to that level.

So yeah, you can be pretty sure that if this track (no pun intended), if this solution is good enough then it will quickly evolve into something that will replace some music creators.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#360
post #247

Earlier quoted context omitted.

As one of the meatsacks whose job you're about to kill... eh, I got nothin, it's damn impressive. It's gonna hit electronic music like a nuclear bomb, I'd wager.

These are tools. Don't think of them as replacements, they aren't. But as tools that will help us be creative. As smart as these apps seem, they will still need a human to decide where and how to use them. They won't replace us but we need to adapt to a new reality.

Kurt Vonnegut and Player Piano has a message for you.
Post reply on HN