Live data from Hacker News

Stable-Audio-Demo

stability-ai.github.io

231–240 of 249 posts

Re: Stable-Audio-Demo

#231
post #92

Earlier quoted context omitted.

I tried generating music on stableaudio.com and, yes, it's bad. However, given the blistering pace of developing in these models, I would not be surprised if these sound incredible in a year or two.

Everyone every time seems to assume a linear (or exponential) curve upwards. But what is the proof for that? I consider it far more likely that we had a breakthrough and now rushing towards the next plateau. Maybe are nearing that. Like in the curve of a PID controller. It's how most or many human improvements go.

I think the proof is seeing how good diffusion models have gotten for making images. They're not perfect but they're leaps and bounds over what we had just a year and a half ago.

Many of these problems seem to have been unexploited simply on basis of nobody throwing enough gpu clusters at it yet.

Re: Stable-Audio-Demo

#232
post #97

Earlier quoted context omitted.

That’s impressive. Why do the printed lyrics for the second chorus differ from the audio? (Which repeats those from the first chorus)

I generated the lyrics using ChatGPT 4 and the suno model attempts to follow them. It generally does a good job, but I have noticed it's fairly common in a second chorus for it to ignore the direction and instead use the same lyrics as the first chorus

That’s fascinating, thanks for clarifying.

Re: Stable-Audio-Demo

#234
post #220
post #203

Earlier quoted context omitted.

Agreed. Leadership can sometimes bring actual value ;) And to be clear, I’m not sure Ed would call himself that. Those are my words, not his.

Ed here. Saw this thread and thought I'd weigh in. Agreed, I wouldn't say I was hired to build Stable Audio. Crazy talented team of research engineers / software engineers / designers did the building. Also wanted to clarify that I didn't quit due to concerns around the training data used for Stable Audio. I was proud of the approach we took to training data - a rev share with rights holders. I quit because of the pr…

FWIW, I am a rightsholder for a number of published songs and recordings. I once spent $12k of my own money on a record and made about $1000 back.

I have spent more blood, tears and money on art than most of you would find even remotely bearable.

I not only consider my songs to be fair use for training a model but I would also honored if my works were included and influenced further musicians in a way that my records probably never will.

The best songwriters I know have other careers and keep on going otherwise. If you actually care about musicians you should make it a habit to go see local live music!

Re: Stable-Audio-Demo

#235

Earlier quoted context omitted.

It isn't unusual for those in leadership positions to use such phrasing when talking about projects and products. It's not a "taking credit" from the engineers sort of thing, but rather about the leadership of the engineers.

Managing a group of people is not synonymous with doing the actual knowledge work of researching and developing innovations that enabled this technology. I find it hard to believe that the contribution of his management somehow uniquely enabled this group of engineers to create this using their experience and expertise. A captain may steer the ship, but they're not the one actually creating and maintaining the means…

> A captain may steer the ship, but they're not the one actually creating and maintaining the means by which it moves.

And yet virtually everyone will go along with a statement like "The captain sailed the ship across the ocean" or "Captain Kirk charted the Gamma Quadrant" or whatever, so I'm not sure how this serves as an objection to the original phrasing.

Re: Stable-Audio-Demo

#236

Not trying to knock the progress here, impressive. As a drummer, 'drum solo' is about as boring as it gets and some weird interspersing sounds. So, it depends on the intended audience. FWIW the sound effects also are not 'realistic' to my ear, at the moment. But again, the progress is huge, well done!

I agree. It's an impressive effort but it's still very far from being able to generate viable music/sound.

There are already millions of library music tracks and sound effects available which sound a lot better. It's going to take a huge investment in gen AI to compete with that and I don't think it makes economic sense (unlike text or images).

Re: Stable-Audio-Demo

#237
post #220
post #203

Earlier quoted context omitted.

Agreed. Leadership can sometimes bring actual value ;) And to be clear, I’m not sure Ed would call himself that. Those are my words, not his.

Ed here. Saw this thread and thought I'd weigh in. Agreed, I wouldn't say I was hired to build Stable Audio. Crazy talented team of research engineers / software engineers / designers did the building. Also wanted to clarify that I didn't quit due to concerns around the training data used for Stable Audio. I was proud of the approach we took to training data - a rev share with rights holders. I quit because of the pr…

Thanks for the clarification, Ed. That’s quite interesting.

Also congrats on the new company!

Re: Stable-Audio-Demo

#238
Why are AI developers so goddamned keen on having it make art, one of the few kinds of work that human beings actually LIKE doing? We could use AI to be a CPA, or to write citations for a paper, but noooo, AI has to be a painter and a musician.

It's almost like the software developers are jealous that someone out there is having a good time and want to take it from them.

Also miss me with that 'AI enables me (a scrub) to make art I couldn't otherwise because I don't want to learn how to do it'. You are lazy. Congrats on finding a high horse about your laziness.

Re: Stable-Audio-Demo

#239

Why are AI developers so goddamned keen on having it make art, one of the few kinds of work that human beings actually LIKE doing? We could use AI to be a CPA, or to write citations for a paper, but noooo, AI has to be a painter and a musician. It's almost like the software developers are jealous that someone out there is having a good time and want to take it from them. Also miss me with that 'AI enables me (a scrub…

I think the development of Generative models for images and audio has more to do with the fact that Computer Vision research goes back decades, and the same systems that originally recognized and labeled images or audio were tweaked to invert the process - and it became naturally an intriguing topic of development precisely because creation is seen as an innately human thing. Beyond that, I'd speculate that the reason we keep seeing developments in "the arts" (though I disagree that an AI can make art, even if it can make beautiful images or music) is because there's no readily-agreed-upon value for that task.

An AI CPA has a specific economic value, but is also a commodity service that no one wants unless they need it. Since there's a clearly comparable cost for needed CPA services, then naturally creating an AI system to do it has a readily comparable market price. People aren't going to make that AI system unless they can do it in way that will make be an improvement as compared to that existing service and price.

I think "just because" has always been a justifiable reason for humans creating beauty (not the same as making art), so it works for research projects better than building a better mousetrap.

Re: Stable-Audio-Demo

#240

Why are AI developers so goddamned keen on having it make art, one of the few kinds of work that human beings actually LIKE doing? We could use AI to be a CPA, or to write citations for a paper, but noooo, AI has to be a painter and a musician. It's almost like the software developers are jealous that someone out there is having a good time and want to take it from them. Also miss me with that 'AI enables me (a scrub…

I think the development of Generative models for images and audio has more to do with the fact that Computer Vision research goes back decades, and the same systems that originally recognized and labeled images or audio were tweaked to invert the process - and it became naturally an intriguing topic of development precisely because creation is seen as an innately human thing. Beyond that, I'd speculate that the reaso…

Thanks for the thoughtful reply! You've given me some stuff to think about
Post reply on HN