Live data from Hacker News

Stable-Audio-Demo

stability-ai.github.io

41–50 of 249 posts

Re: Stable-Audio-Demo

#41
post #35

Interestingly, Ed Newton-Rex, the person hired to build Stable Audio, quit shortly after it was released due to concerns around copyright and the training data being used. He’s since founded https://www.fairlytrained.org/ Reference: https://x.com/ednewtonrex

For generative models, if the model authors do not publish the architecture of their model; and, the model uses a transformation from text to another kind of media; you can assume that they have delegated some part of their model to a text encoder or similar feature which is trained on data that they do not have an express license to.

Even for rightsholders with tens of millions to hundreds of millions of library items like images or audio snippets, the performance of the encoder or similar feature in text-to-X generative models is too poor on the less than billion tokens of text in the large repositories. This includes Adobe's Firefly.

It is also a misconception that large amounts of similar data, like the kinds that appear in these libraries, is especially useful. Without a powerful text encoder, the net result is that most text-to-X models create things that look or sound very average.

The simplest way to dispel such issues is to publish the architecture of the model.

But anyway, even if it were all true, the only reason we are talking about diffusers, and the only reason we are paying attention to this author's work Fairly Trained, is because of someone training on data that was not expressly licensed.

Re: Stable-Audio-Demo

#43
post #36
post #26

I think we still need the step where the AI learns what a high quality sound library sounds like and then applies the previously learned abilities by triggering sounds of that library via MIDI. That way you'd get perfect audio quality with the creativity of a musical AI.

How would MIDI get you eg a guitar being played dirty? Or some subtle echo that comes from recording in a bathroom?

It would use a sampler and for the subtle echo effect add a reverb to the bus.

https://www.youtube.com/watch?v=EQdp2QLiSYQ&t=187s

Re: Stable-Audio-Demo

#44
post #35

Interestingly, Ed Newton-Rex, the person hired to build Stable Audio, quit shortly after it was released due to concerns around copyright and the training data being used. He’s since founded https://www.fairlytrained.org/ Reference: https://x.com/ednewtonrex

For generative models, if the model authors do not publish the architecture of their model; and, the model uses a transformation from text to another kind of media; you can assume that they have delegated some part of their model to a text encoder or similar feature which is trained on data that they do not have an express license to. Even for rightsholders with tens of millions to hundreds of millions of library ite…

If you require licensing fees for training data, you kill open source ML.

That’s why it’s important for OpenAI to win the upcoming court cases.

If they lose, they’ll survive. But it will be the end of open model releases.

To be clear, I don’t like the idea of companies profiting off of people’s work. I just like open source dying even less.

Re: Stable-Audio-Demo

#45

This is right into the "uncanny valley" of music. It definitely sounded "like music", but none of it is what a human would produce. There's just something off.

Here is a silly song I generated using suno.ai, which I have found to be incredibly impressive (at least, a small percentage of its outputs are very good, most are bad). I think it's good enough that most humans wouldn't realise it's AI generated. https://app.suno.ai/song/8a64868d-9dd3-46db-91af-f962d4bec8b...

Re: Stable-Audio-Demo

#46

Earlier quoted context omitted.

For generative models, if the model authors do not publish the architecture of their model; and, the model uses a transformation from text to another kind of media; you can assume that they have delegated some part of their model to a text encoder or similar feature which is trained on data that they do not have an express license to. Even for rightsholders with tens of millions to hundreds of millions of library ite…

If you require licensing fees for training data, you kill open source ML. That’s why it’s important for OpenAI to win the upcoming court cases. If they lose, they’ll survive. But it will be the end of open model releases. To be clear, I don’t like the idea of companies profiting off of people’s work. I just like open source dying even less.

[deleted]

Re: Stable-Audio-Demo

#47

Earlier quoted context omitted.

For generative models, if the model authors do not publish the architecture of their model; and, the model uses a transformation from text to another kind of media; you can assume that they have delegated some part of their model to a text encoder or similar feature which is trained on data that they do not have an express license to. Even for rightsholders with tens of millions to hundreds of millions of library ite…

If you require licensing fees for training data, you kill open source ML. That’s why it’s important for OpenAI to win the upcoming court cases. If they lose, they’ll survive. But it will be the end of open model releases. To be clear, I don’t like the idea of companies profiting off of people’s work. I just like open source dying even less.

Replying to a deleted comment:

> It sounds as if you imply that would be bad. But what if it wasn't?

Entirely possible. The early history of aviation was open source in the sense that many unlicensed people participated, and died. The world is strictly better with licensing requirements in place for that field.

But no one knows. And if history is any guide for software, it seems better to err on freedoms that happen to have some downside rather then clamping down on them. One could imagine a world where BitTorrent was illegal. Or cryptography, or bitcoin.

Re: Stable-Audio-Demo

#48
> We append “high-quality, stereo” to our sound effects prompts because it is generally helpful.

It's hilarious that we've discovered you can get better outputs from LLMs by simply nicely telling it to generate better results.

Re: Stable-Audio-Demo

#49
post #26

I think we still need the step where the AI learns what a high quality sound library sounds like and then applies the previously learned abilities by triggering sounds of that library via MIDI. That way you'd get perfect audio quality with the creativity of a musical AI.

I've always wished for something like that for image generation AI. It'd be much cooler/more interesting to watch AI try to draw/paint pictures with strokes rather than just magically iterate into a fully-rendered image. I dunno what kind of dataset or architecture you could possibly apply to accomplish this, but it would be very interesting.
Post reply on HN