Live data from Hacker News

Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

stability.ai

141–150 of 210 posts

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#141
post #107

Earlier quoted context omitted.

lol, no. Not autogenerate midi (although their latest versions of BiaB are pretty darn good now) but generate waveforms together. It would be similar to having AI generate whole scores of music but ensuring it's all in sync and in key. Not taking sample database of 88 sound files and triggering them when the midi-note strikes.

That's not how BiaB works at all. It has all kinds of patterns built into it. So, it knows, how to generate, say, a bluegrass bassline in a given key. There are plenty of ways to play back MIDI with high sound quality, including feeding it into an AI-driven VST like NotePerformer.

Then explain why you categorize it as cheesy? Sounds like it’s pretty cool.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#142

The solo piano was interesting because of how clean it is. I can imagine going from that sample to a score without too much difficulty. Once it's in a symbolic format it becomes much more flexible and re-usable. While this does not seem to be the trend I hope more gen ai in the audio and visual realms start to produce more structured / symbolic output. For example, if I were Adobe I would be training models, not to o…

I strongly agree about generating "editables" rather than finalized media. In fact, that's why text generators are more useful than current media generators: text is editable by default. Here's a tweetstorm about it: https://x.com/jsonriggs/status/1694490308220964999?s=20

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#143
post #107

Earlier quoted context omitted.

That's not how BiaB works at all. It has all kinds of patterns built into it. So, it knows, how to generate, say, a bluegrass bassline in a given key. There are plenty of ways to play back MIDI with high sound quality, including feeding it into an AI-driven VST like NotePerformer.

Then explain why you categorize it as cheesy? Sounds like it’s pretty cool.

he's saying the output is cheesy. it sounds like stuff you'd hear on a demo track for a kid's toy piano

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#144
post #21

Earlier quoted context omitted.

Why not? Mostly for private use in my case. SDXL has created some beautiful works of art in my experiments and I would love to have a similar experience in the music world.

Come on, be creative and make something new instead of copying someone else. It’s just kind of lame imo. “Mostly” private use? Mmm. :thumbs down emoji:

there is very little creativity in most music already

(Axis of Awesome - 4 Four Chord Song)

https://youtu.be/5pidokakU4I?t=52

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#145

The solo piano was interesting because of how clean it is. I can imagine going from that sample to a score without too much difficulty. Once it's in a symbolic format it becomes much more flexible and re-usable. While this does not seem to be the trend I hope more gen ai in the audio and visual realms start to produce more structured / symbolic output. For example, if I were Adobe I would be training models, not to o…

I strongly agree about generating "editables" rather than finalized media. In fact, that's why text generators are more useful than current media generators: text is editable by default. Here's a tweetstorm about it: https://x.com/jsonriggs/status/1694490308220964999?s=20

Audio is definitely editable. While generative audio is new I am hopeful that a host of interesting applications will emerge (audio2audio etc.) within its ecosystem. Promising signal separation (audio to STEMs) and pitch detection tools already exist for raw audio signals. If you want to force Stability to focus on symbolic representations (such as severely lossy MIDI) I hope you can instead first try adapting to tools that work fundamentally with rich audio signals. Perhaps there will be room for symbolic music AI and perhaps Stability will even develop additional models that generate schematic music, but please please don't sacrifice audio generality for piano roll thinking alone. LORAs will undoubtedly be usable to generate more schematic audio via the Stable Audio model -- I imagine they could be easily purposedly to develop sample libraries compatible with DAW (digital audio workstation), sequencer and tracker production workflows.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#146

Earlier quoted context omitted.

You joke, but even those videos of people saying their dog can talk is just like this. It's cute because it's a real dog making sounds we really want to believe when it's just them mimicking sounds because they get pettin's and treats. What I want is "AI" to do something impressive. Why are we trying to make the system generate the sounds itself? We don't make artists do that, we give them instruments. Give the model…

We work with what we have. We don't have a lot of recordings of the physical movements of musicians; we have recordings. Similarly, we don't have recordings of the actions of painters; we have finished paintings -- but if you're not impressed with what AI can do in the visual sphere, your standards are, to put it mildly, high.

I'm not really sure how to take this. We absolutely have recordings of instruments. You can buy them as complete sets. You train on complete recordings, and then tell it how to use the sampled instruments to compose a song in the style of the trained data. Building something to make a waveform that looks like another waveform just seems like a very odd direction to take.

Yes, my standards are if it isn't at least as good as what's available now, what's the point.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#147
post #6

The bluegrass one is super weird. I can’t identify exactly why.

I can identify a bunch of things. The chord structure jumps all over randomly in a genre that usually does the opposite. The banjo is clearly not an actual banjo being strummed/frailed, but a weird agglomeration of bright toned instruments including both frailed/scruggs banjo and dobro, and maybe harmonica and fiddle creeping in. The AI doesn't know it's making a combination of instruments, so where it's trained on instruments blending, it thinks it can produce pre-blended sounds. I guess maybe this is more like a return to being a child hearing music for the first time with no preconceptions or expectations.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#148

Earlier quoted context omitted.

Build or seek out a MIDI generating model. I hope Stable Audio is never the place for that. MIDI is deeply lossy and it would be tragedy if it was the only music representation. Imagine if instead of phonographs, compact disks and streaming audio we only had piano rolls. What a loss indeed.

Midi is not lossy, midi is symbolic. There's a huge difference.

MIDI is an extraordinarily lossy music representation. Even Claude Shannon would facepalm at the assertion that it could, in theory, represent audio faithfully. It is not its purpose, it is decidedly not its practice, and it is a ludicrously irrelevant example of pedantry to say otherwise. The false equivalency asserted by the commons can be aggravating :D

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#149

Earlier quoted context omitted.

No, it's lossy. It's an event model at a fixed data rate. You can only do so many things sequentially, even if you could represent any possible musical concept as a MIDI event. So even if you're not sticking to note-on, note-off, it's still extremely lossy.

MIDI is able to accommodate nearly everything that can be represented through a musical score and instrumental performance. What are you hoping to accomplish with AI-generated waveforms that can't be done with MIDI?

The problems have long been known and articulated: http://www.music.mcgill.ca/~gary/courses/papers/Moore-Dysfun...

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#150

Earlier quoted context omitted.

No, it's lossy. It's an event model at a fixed data rate. You can only do so many things sequentially, even if you could represent any possible musical concept as a MIDI event. So even if you're not sticking to note-on, note-off, it's still extremely lossy.

MIDI is able to accommodate nearly everything that can be represented through a musical score and instrumental performance. What are you hoping to accomplish with AI-generated waveforms that can't be done with MIDI?

> What

"Intention" (as a tentative term)

The question becomes: what has impeded the creation of a MIDI file that can be confused with an actual concert from Arturo Benedetti Michelangeli.

Post reply on HN