Live data from Hacker News

Riffusion – Stable Diffusion fine-tuned to generate music

riffusion.com

281–290 of 481 posts

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#283

I read the article: "If you have a GPU powerful enough to generate stable diffusion results in under five seconds, you can run the experience locally using our test flask server." Curious what sort of GPU the author was using or what some of the min requirements might be?

Author here: fwiw we are running the app on a10g GPUs, which generally can turn around a 512x512 in 3.5s with 50 inference steps. This time includes converting the image into audio which should be done on the GPU as well for real-time purposes. We did some optimization such as a traced unet, fp16 and removing autocast. There are lots of ways it could be sped up further I'm sure!

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#284

personalized RL agents that finds aesthetic trajectories through the music latent space... soon, i hope :D

Love this idea. If I had more time I wanted to make a spaceship game where you are flying around the latent space, and model interrogation is used to provide labels to landmarks as you move around.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#285
post #196
post #188

Earlier quoted context omitted.

Does the institute not also claim that work?

Fair enough. But for me, when talking about `having an insight`, I don't imagine a non-human entity doing that. And to be pedantic (talking about Germans doing research, I hope everyone would expect me to be), the institute is called Fraunhofer IIS. `Fraunhofer` would colloquially refer to the society, which is an organization with 76 institutes total. Although, of course, the society will also claim the work...

It's an interesting question, one I hadn't thought of before. But in common language, it sometimes makes sense to credit the institution, others just the individuals. I think may be more based around how much the institution collectively presents itself as the author and speaks on behalf of the project versus the individuals involved. Here is my own general intuition for a few contrasting cases:

Random forests: Ko and Breiman's, not really Bell Labs and UC-Berkeley

Transistors: Bardeen, Brattain, and Shockley, not really Bell Labs (thank the Nobel Prize for that)

UNIX: Primarily Bell Labs, but also Ken Thompson and Dennis Richie (this is a hard one)

GPT-n: OpenAI, not really any individual, and I can't seem to even recall any named individual from memory

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#286

Earlier this year, graphic designers, last month it was software engineers, and now musicians are also feeling the effects. Who else will AI make looking for a new job?

If I was a musician, this post would not make me worry for a second

If a hack based on an image generator already has promising results for music generation, then imagine what will happen if something dedicated to music is built from the ground up.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#288

Earlier quoted context omitted.

This was the first AI thing to fill me with a feeling of existential dread.

What is with the hyperbole in this thread? This stuff sounds like incoherent noise. It is noticeably worse than AI audio stuff I heard 5 years ago. What is going on with the responses here?

Usage of an image generator to produce passable music fragments, even if they sound a bit distorted, is very surprising. That type of novelty is why we come here.

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#289
Multiple folks have asked here and in other forums but I'm going to reiterate, what data set of paired music-captions was this trained on? It seems strange to put up a splashy demo and repo with model checkpoints but not explain where the model came from... is there something fishy going on?

Re: Riffusion – Stable Diffusion fine-tuned to generate music

#290
post #260

Earlier quoted context omitted.

In general all this stuff is chopping the bottom off the market. AI art, code, writing, music, etc. can all generate passable "filler" content, which will decimate all human employment generating same. I don't think this stuff is a threat to genuinely innovative, thoughtful, meaningful work, but that's the top of the market. That being said the bottom of the market is how a lot of artists make their living, so this i…

The only thing that affects whether you have a job is the Federal Reserve, not how good productivity tools are. You always have comparative advantage vs an AI, so you always have the qualifications for an entry level job. There will never be a revolution and there's no such thing as late capitalism. Well, not if the Fed does their job.

I see a lot of AI naysayers neglecting the comparative advantage part.

If AI completely eliminates low skill art labour from the job pool, it's not like those affected by it are gonna disintegrate, riot, and restructure society. They have the choice of filling an art niche an AI can't or they can spend that time learning other, more in-demand skills. This also ignores that fact that some companies would rather reallocate you to more profitable projects even if your art skills don't change.

Selling a product with relative value like a painting or a sculpture will always be an uphill battle. Now that there's more competition from AI, it just gives artists/businesses incentive to find what people want that an AI can't deliver. Worst case scenario, employment rates in this sector are rough while the market recalibrates. Interested to see how these technologies develop.

Post reply on HN