Live data from Hacker News

DiffRhythm: Fast End-to-End Full-Length Song Generation with Latent Diffusion

aslp-lab.github.io

71–80 of 121 posts

Re: DiffRhythm: Fast End-to-End Full-Length Song Generation with Latent Diffusion

#73

If I am to retain any interest as an amateur music writer without proaudio engineering skills and equipment, but with a day job, , I want tools that help me enact MY vision to reality. That means multi tracking, ability to hum or score a melody and have it transfer to musical instrument, ability to enter existing tracks, provide a temporal segment for diffusion, and ask it to 'generate a counterpoint to the melody wi…

I definitely see this happening. Music generation has lagged behind image generation but is following more or less the same path. Early image generation models were completely unconditional; all you could do was sample an image. Then coarse conditioning methods such as text prompts and depth images came along; then additional tooling to tune images in a more fine-grained way.

That said, there is a difference to images in that music also has a "symbolic" level to it that is closer to text than images [1]. There's other work out there that uses LLM-type tools for direct melody generation (no audio). And of course, there's lyrics. I do expect commercial tools to start integrating all these capabilities gradually, it's just a matter of time.

[1] I guess there's also vector images (like SVG) - I've seen work in generating those as well, though it's less mature than directly generating pixels.

Re: DiffRhythm: Fast End-to-End Full-Length Song Generation with Latent Diffusion

#74

One thing that strikes me about almost every AI-generated track (from academic or commercial generators), is that even if it's often "competent" - in that it has reasonable melodies, chord progressions, etc - is how average it is. Mediocre, taking the term literally. In a way that also highlights cliches and crutches that are common in human-made music. Somewhat reminiscent of GPT text that drones on and on in a gram…

Yeah-- in a professional workflow, at best, these tools are for getting ideas rather than creating output that will be used directly. Lots of folks use them for actual creation because they're just so enamored with the ability to create vaguely technically competent output from text, but they're all pretty much a bee-line to mediocre, and overcoming mediocrity is absolutely the most difficult part of working with AI output. The same is true with text, as you mentioned, and image generators. As Charles Eames said, "The details are not the details. They make the design." Well, these tools suck with details, and details convey character, perspective, message, meaning, etc. Surely the tooling will improve this in years to come, but it certainly hasn't yet.

Re: DiffRhythm: Fast End-to-End Full-Length Song Generation with Latent Diffusion

#75

Cool. Obviously needs some work. Lots of artifacts. Something to build on though. Lots of sour grapes comments from folks. Too bad. Not what I expect out of Hacker News. Glad people are pushing the technological envelope and exploring this space despite the strong negative emotions.

[flagged]

Re: DiffRhythm: Fast End-to-End Full-Length Song Generation with Latent Diffusion

#76
post #65

Earlier quoted context omitted.

What an incredibly elitist, smug attitude. You're basically saying people only have the right to hear the music that professionals think they should hear.

That's not smug at all. That's not what I'm saying either. It's just that.. you can't master something you don't practice and understand. It's true in every single thing in life you do, sports, literature, maths, music, cuisine, kindness, etc. If you don't like to compose music, why suffer this and even submit to the randomness of some computer program, rather than giving the opportunity to another fellow human to op…

I don't understand the "give it [the task?] to someone who does" part. Obtaining a hobbyist composer who is available at short notice and obeys instructions for free is not usually an option. Maybe there's a website for this, but it would have to be humming with idle composers in order to offer quick and satisfactory results.

I think "stop not enjoying it" is a better line to take. Like with AI illustrations (where I'd much rather see a blog author's crappy biro drawings instead), terrible amateur efforts with some online 808 emulator or whatever would be more entertaining and interesting than AI output.

Re: DiffRhythm: Fast End-to-End Full-Length Song Generation with Latent Diffusion

#77

Business hates creatives. They'll do anything to automate us away.

If I was a business I'd "hate" creatives too, and I'd also want to automate them away. The costs of producing (truly) creative works is utterly bonkers, and so are the risks associated.

That's why corporations that have made creative products have traditionally never gone anywhere. They all just went out of business. And all the artists got rich.

Re: DiffRhythm: Fast End-to-End Full-Length Song Generation with Latent Diffusion

#78

[flagged]

>You are automating an activity humans ENJOY doing. There's at least an order of magnitude more people who enjoy making music than there are people with the actual skill/talent to make music. Music generation AI is an absolute blessing to the untalented among us who'd love to make a song in a certain style or with certain lyrics but lack the time, talent or ability to do it ourselves.

Good for them, really.

But don't mistake one thing for the other: how is it different than, say, being Emperor Joseph II asking Mozart in Vienna to write an opera for him?

Mozart wrote the music, not Joseph.

Similarly, you can hike across France, from South to Britain for several days. Or you can take the train. Or a car, alone, or with a driver. Or a plane, in the pilot or the passenger seat.

You'll get in the same place in the end. The experience will be totally, fundamentally different for you, as well as for others.

Re: DiffRhythm: Fast End-to-End Full-Length Song Generation with Latent Diffusion

#79

It’s just combining sample WAV files without human coordination, talk about a lame-ass achievement. It’s already easy enough to set BPM and load in files in Ableton and warp them into unison, from what I heard this is basically just that with”HOORAY FOR AI” slathered as a veneer on top. If you think I’m being harsh, I have my reasons as a professional musician to critique these things in an unflattering light because…

Sorry no. Here on HN, your having a vested interest in some market makes your opinion entirely invalid. That is, enless you're interested in one of the correct markets such as software or AI services.

Re: DiffRhythm: Fast End-to-End Full-Length Song Generation with Latent Diffusion

#80

Earlier quoted context omitted.

If I was a business I'd "hate" creatives too, and I'd also want to automate them away. The costs of producing (truly) creative works is utterly bonkers, and so are the risks associated.

That's why corporations that have made creative products have traditionally never gone anywhere. They all just went out of business. And all the artists got rich.

?? How do you think what you say follows from what I said?
Post reply on HN