Live data from Hacker News

Riffusion – Stable Diffusion fine-tuned to generate music

riffusion.com

461–470 of 481 posts

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#461

Earlier quoted context omitted.

Have you tried AI asset generators? They are working extremely good. Just yesterday a friend of mine has shown me the progress they made in their game. It is incredible. Designers are 100% loosing their job over this.

I'm a professional game developer and excited AI enthusiast. While I've seen a lot of cool stuff which helps generation for hobby projects or smaller indie games it's nowhere near the quality and consistency needed to come close to the work of a skilled human artist at a larger studio.

Yes, and I studied Game Development in Germany for 3 years in Düsseldorf, got third at the national gameforge newcomer award with my team "Northlight Games" and still have many connections to the people in the business (if this somehow matters). The quality for 2d assets is on very professional level and already replaced jobs in projects I know.

To give you an example join the public discord of https://www.scenario.gg and check out results. Come back and tell me those aren't on a professional level.

I am not saying that designers won't be needed anymore but AI is definitely able to replace jobs and speed up progress in game development.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#462

Earlier quoted context omitted.

8GB is enough to do 1080p resolution. the UI i use for SD maxes out at 2048x2048. however, it takes a lot longer than 512x512 to generate: 1m40s versus 1.97s. I'm guessing if one had access to one of those nvidia backplane rackmount devices one could generate 8k or larger resolution images.

SD can’t generate coherent images if you increase the output size. They’re basically always unusable unless you don’t need any global architecture to them.

I'm not sure what you mean. with the 1.4 checkpoint i can make a scene "a sky full of blimps" and then img2img that into a space battle where all the blimps in the sky are either fires or spacecraft afterburners. at 2k pixels, and then use the built in "GAN" stuff to bring it to 4k, 5k, or 8k.

It isn't perfect and it takes a lot of fiddling and sysadmin stuff, but it is way better than nextchar = CHR$(RAND64); color = (rand65535); of yore.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#463

Earlier quoted context omitted.

So what is phase? From dabbling with waveforms in audio editors, sampling, and later learning a little bit about complex numbers, phase seems eventually equivalent to what would sound like changing pitch, modulating the frequency of a periodic signal. The simplest demonstration of it is the doppler shift. But it's not at all that simple because moving relative to the source the sound pressure and thus the perceived l…

Phase is the offset in time. The functions sin(θ) and sin(θ + c), for arbitrary real c, represent the same frequency signal; they are offset from each other horizontally by c, and that c is a phase difference. It has an interpretation as an angle, when the full cycle of the wave is regarded as degrees around a circle; and that's what I mean by rotating phase. When you take a window of samples of a signal, and run the…

> on the hypothesis that we have a pure, periodic signal in there

That pure sine wouldn't generate any artefacts. It would result in a 200Hz output from the AI if it throws the phase information out. You wouldn't hear a difference unless its an (aptly so called) complex signal. Eg. 200 and 201 Hz layered is an impure signal with a period below 1Hz, far outside the scope. Eventually the signals will cancel out completely. [1]

The important point is, I think, that FFT doesn't simply look at the offset aka phase. Rather, 201 Hz looks like a 200 Hz that is moving. So it encodes phase-shift in the delta of the offset between two windows. For a sum of 200 and 201 Hz it has to assume that the magnitude is also changing, which I find entirely counterintuitive.

From the mathematical perspective, this seems like a borring homework, far detached from accoustics. So, I don't know. The funny thing is that rotation is very real in the movement of strings. If the orbit in one point is elliptic, that's like two sinusoids at different magnitudes offset by some 90 degree, in a simplified model. But it has nearly infinite coupled points along its axis. As they exite each other, and each point has a different distance to the receiver, that's where phase shift happens.

> If you look here at the definition of the ω (omega) parameter

I wasn't going to make drone, but I will take a look.

1: https://graphtoy.com/?f1(x,t)=100*sin(x)&v1=true&f2(x,t)=100...

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#464

Is there a different mapping of FFT information to a two dimensional image that would make harmonic relationships more obvious? IE, use a polar coordinate system where angle 12 oclock is 440hz, and the 12 chromatic notes would be mapped to the angle of the hours. Maybe red pixel intensity is bit mapped to octave, IE first, third and eight octave: 0b10100001. Time would be represented by radius. Unfortunately the spac…

You need the mapping to be only to 1 dimension, since you need the 2nd dimension for time

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#465

This opens up ideas. One thing people have tried to do with stable diffusion is create animations. Of course, they all come out pretty janky and gross, you can't get the animation smooth. But what if what if a model was trained not on single images, but animated sequential frames, in sets, laid out on a single visual plane. So a panel might show a short sequence of a disney princess expressing a particular emotion as…

a different approach: analog television

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#466

Was just watching an interview of Billy Corgan (smashing pumpkins) on Rick Beato’s YouTube[1] last night where billy was lamenting the inevitable future where the “psychopaths” in the music biz will use ai and auto tune to churn out three chord non-music mumble rap for the youth of tomorrow, or something to that effect. It was funny because it’s the sad truth. It’s already here but new tech will allow them to cut cos…

> new tech will allow them to cut costs even more, and increase their margins How, when everyone and their dog can generate such music? It's gonna be like stock photography in the age of SD.

They do marketing not music.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#467
post #407

If copyright laws don't catch up, the sampling industry is cooked. Made this: https://soundcloud.com/obnmusic/ai-sampling-riffusion-waves-...

assuming the first sample was generated by OP method, how did you clean that sample up so nicely?

Yep, the first sample was generated via OP's method. The tl;dr answer is that I used a few plugins in my digital audio workstation to make it sound better.

I made a video specifically in reply to you if you want to see exactly how I did it (3 mins): https://www.youtube.com/watch?v=69Q-cseNCI4

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#468

Other author here! This got a posted a little earlier than we intended so we didn't have our GPUs scaled up yet. Please hang on and try throughout the day! Meanwhile, please read our about page http://riffusion.com/about It’s all open source and the code lives at https://github.com/hmartiro/riffusion-app --> if you have a GPU you can run it yourself This has been our hobby project for the past few months. Seeing the…

Amazing work! Do you plan on open-sourcing the code to train the model?

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#469
post #247

Earlier quoted context omitted.

As one of the meatsacks whose job you're about to kill... eh, I got nothin, it's damn impressive. It's gonna hit electronic music like a nuclear bomb, I'd wager.

As a listener, I think you're probably still safe. Can you use this to help you though? Maybe. It's impressive what it produces, but I think it probably lacks substance in the same way the visual AI art stuff does. For the most part, it passes what I call the at-a-glanceness test. It's little better than apophenia (the same thing that makes you see shapes in clouds, faces in rocks, or think you've recognised a famili…

I think this is a good point. To make this useful for music creators, and to make music creation more generally accessible, the output needs to be more useful. We are working on that at https://neptunely.com

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#470

This is awesome! It would be interesting to generate individual stems for each instrument, or even MIDI notes to be rendered by a DAW and VST plugins. It's unfortunate that most musicians don't release the source files for their songs so it's hard to get enough training data. There's a lot of MIDI files out there but they don't usually have information about effects, EQ, mastering, etc.

Let's get in touch. This is precisely what we are working on at Neptunely. https://neptunely.com
Post reply on HN