Live data from Hacker News

Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

stability.ai

11–20 of 210 posts

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#13
post #7

Everything looks very convincing apart from the airpane pilot and the sound effects. They sound very weird as if one is hallucinating

The airplane one just sounds like a foreign language over a bad intercom, I think that could still be useful for some stuff.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#14

Now imagine Spotify using this to generate individual earworms for everybody based on their personal tastes (likes, playlists). Yes, AI is partly hype, but had someone told me this even two years ago, I wouldn't have believed it.

This is why it’s vital that AI is openly available. Imagine a world where Spotify is the only company that can do that, and they use it to make sure they never pay royalties again.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#15

Earlier quoted context omitted.

We've been researching an audio denoiser for music that we will present at the AES conference in October. Description page: https://tape.it/denoising We'll also publish a webapp where you can use the denoiser for free. Mail me if you want beta access to it (email in profile). It won't be open-source though, although the paper will of course be public. It will also only reduce noise, and not reconstruct other aspects…

denoising seems to fail in the guitar and vocals example

Can you clarify where it fails? It's designed to remove stationary noise only, and removes it very well in the guitar and vocals example.

Generally speaking, if you have other sounds that you don't want in the audio, we don't remove them - it's hard to decide from a musical point of view whether you want a certain sound or not. To give an extreme example: a barking dog probably doesn't belong into a Zoom conference, but it may very well belong into your audio recording. Removing such elements would be a creative decision.

The guitar and vocals example has certain clicks in the background that we don't remove - but the stationary noise is gone. Existing professional (and complex) audio restoration tools like iZotope RX don't remove those clicks, either. It's a conservative approach, sure, but in return you can throw any audio at it and it always improves it.

Re: Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion

#16
post #3

Thank you for sharing! On a tangent: I'm wondering if there are any good open source models/libraries to reconstruct audio quality. I'm thinking about an end-to-end open source alternative to something like Adobe Podcast [1] to make noisy recordings sound professional. Anecdotally it's supposed to be very good. In a recent search, I haven't found anything convincing. In my naive view this tasks seems much simpler tha…

There seems to have been a fork in the road:

On one side the tech for literal denoising has stagnated a bit. It’s a very hard problem to remove all noise while keeping things like transients.

On the other side, AI is being rapidly developed for it’s ability to denoise by recreating the recording, just without the noise.

Post reply on HN