Earlier quoted context omitted.
lol, no. Not autogenerate midi (although their latest versions of BiaB are pretty darn good now) but generate waveforms together. It would be similar to having AI generate whole scores of music but ensuring it's all in sync and in key. Not taking sample database of 88 sound files and triggering them when the midi-note strikes.
That's not how BiaB works at all. It has all kinds of patterns built into it. So, it knows, how to generate, say, a bluegrass bassline in a given key. There are plenty of ways to play back MIDI with high sound quality, including feeding it into an AI-driven VST like NotePerformer.
Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
141–150 of 210 posts
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#142The solo piano was interesting because of how clean it is. I can imagine going from that sample to a score without too much difficulty. Once it's in a symbolic format it becomes much more flexible and re-usable. While this does not seem to be the trend I hope more gen ai in the audio and visual realms start to produce more structured / symbolic output. For example, if I were Adobe I would be training models, not to o…
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#143Earlier quoted context omitted.
That's not how BiaB works at all. It has all kinds of patterns built into it. So, it knows, how to generate, say, a bluegrass bassline in a given key. There are plenty of ways to play back MIDI with high sound quality, including feeding it into an AI-driven VST like NotePerformer.
Then explain why you categorize it as cheesy? Sounds like it’s pretty cool.
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#144Earlier quoted context omitted.
Why not? Mostly for private use in my case. SDXL has created some beautiful works of art in my experiments and I would love to have a similar experience in the music world.
Come on, be creative and make something new instead of copying someone else. It’s just kind of lame imo. “Mostly” private use? Mmm. :thumbs down emoji:
(Axis of Awesome - 4 Four Chord Song)
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#145The solo piano was interesting because of how clean it is. I can imagine going from that sample to a score without too much difficulty. Once it's in a symbolic format it becomes much more flexible and re-usable. While this does not seem to be the trend I hope more gen ai in the audio and visual realms start to produce more structured / symbolic output. For example, if I were Adobe I would be training models, not to o…
I strongly agree about generating "editables" rather than finalized media. In fact, that's why text generators are more useful than current media generators: text is editable by default. Here's a tweetstorm about it: https://x.com/jsonriggs/status/1694490308220964999?s=20
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#146Earlier quoted context omitted.
You joke, but even those videos of people saying their dog can talk is just like this. It's cute because it's a real dog making sounds we really want to believe when it's just them mimicking sounds because they get pettin's and treats. What I want is "AI" to do something impressive. Why are we trying to make the system generate the sounds itself? We don't make artists do that, we give them instruments. Give the model…
We work with what we have. We don't have a lot of recordings of the physical movements of musicians; we have recordings. Similarly, we don't have recordings of the actions of painters; we have finished paintings -- but if you're not impressed with what AI can do in the visual sphere, your standards are, to put it mildly, high.
Yes, my standards are if it isn't at least as good as what's available now, what's the point.
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#147The bluegrass one is super weird. I can’t identify exactly why.
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#148Earlier quoted context omitted.
Build or seek out a MIDI generating model. I hope Stable Audio is never the place for that. MIDI is deeply lossy and it would be tragedy if it was the only music representation. Imagine if instead of phonographs, compact disks and streaming audio we only had piano rolls. What a loss indeed.
Midi is not lossy, midi is symbolic. There's a huge difference.
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#149Earlier quoted context omitted.
No, it's lossy. It's an event model at a fixed data rate. You can only do so many things sequentially, even if you could represent any possible musical concept as a MIDI event. So even if you're not sticking to note-on, note-off, it's still extremely lossy.
MIDI is able to accommodate nearly everything that can be represented through a musical score and instrumental performance. What are you hoping to accomplish with AI-generated waveforms that can't be done with MIDI?
Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
#150Earlier quoted context omitted.
No, it's lossy. It's an event model at a fixed data rate. You can only do so many things sequentially, even if you could represent any possible musical concept as a MIDI event. So even if you're not sticking to note-on, note-off, it's still extremely lossy.
MIDI is able to accommodate nearly everything that can be represented through a musical score and instrumental performance. What are you hoping to accomplish with AI-generated waveforms that can't be done with MIDI?
"Intention" (as a tentative term)
The question becomes: what has impeded the creation of a MIDI file that can be confused with an actual concert from Arturo Benedetti Michelangeli.