Live data from Hacker News

Stable-Audio-Demo

stability-ai.github.io

211–220 of 249 posts

Re: Stable-Audio-Demo

#211

Earlier quoted context omitted.

> If you require licensing fees for training data, you kill open source ML. This is another one of those “well if you treat the people fairly it causes problems” sort of arguments. And: Sorry. If you want to do this you have to figure out how to do it ethically. There are all sorts of situations where research would go much faster if we behaved unethically or illegally. Medicine, for example. Or shooting people in ro…

"Ethical" in this case is a matter of opinion. The whole point of copyright was to promote useful sciences and arts. It’s in the US constitution. You don’t get to control your work out of some sense of fairness, but rather because it’s better for the society you live in. As an ML researcher, no, there’s basically no way to make progress without the data. Not in comparison with billion dollar corporations that can thr…

Yes, it’s true that open source projects that cannot pay to license content owned by other people are at a disadvantage versus those who can. Open source projects cannot, for example, wholly copy code owned by other people.

Also, beware of originalist interpretations of the Constitution. I believe there’s been about 250 years of law clarifying how copyright works, and, not to beat a dead horse, I don’t think it carves out a special exception for open source projects.

Re: Stable-Audio-Demo

#212

This is right into the "uncanny valley" of music. It definitely sounded "like music", but none of it is what a human would produce. There's just something off.

AI pictures are the same. We are more tolerant of six fingered-pictures with missing limbs, for some reason.

We're used to drawings, 3D renders, etc.

There's no such thing as "artificial music" - at the very least, not since electronic music has become mainstream.

Re: Stable-Audio-Demo

#213

Earlier quoted context omitted.

If you require licensing fees for training data, you kill open source ML. That’s why it’s important for OpenAI to win the upcoming court cases. If they lose, they’ll survive. But it will be the end of open model releases. To be clear, I don’t like the idea of companies profiting off of people’s work. I just like open source dying even less.

The reality is always a dynamic tension between law, regulation, precedent, and enforceability. It is possible to strangle OpenAI without strangling AI: pmarca is anti-OpenAI in print, but you can bet your butt he hopes to invest in whatever replaces it, and he’s got access to information that like, 10 people do. A useful example would be the Napster Wars: the music industry had been rent seeking (taking the fucking…

a16z are investors in openai

Re: Stable-Audio-Demo

#214

This is incredibly good compared to SOTA music models (MusicGen, MusicLM). It looks like there's also a product page where you can subscribe to use it, similar to Midjourney: https://www.stableaudio.com/ Sadly it's not open-weight and it doesn't look like there's an API (again like Midjourney): you subscribe monthly to generate audio in their UI, rather than having something developers can integrate or wrap.

I was hoping to use it to generate some sound effects to use in a game I'm working on - but looks like I need an "enterprise license" ( https://www.stableaudio.com/pricing ) Why does this have a different clause I wonder, and doesn't just fall under "In commercial products below 100,000 MAU"?

Different deal with the underlying data holders with revenue share etc

Re: Stable-Audio-Demo

#215
post #184

Earlier quoted context omitted.

stableaudio.com is fully licensed, music is an interesting area https://www.musicbusinessworldwide.com/stability-ai-launches...

Serious question, I'd genuinely like to know - why? You didn't license the images when training Stable Diffusion, and yet you did for Stable Audio? In both cases the training should either be fair use and legal without any licensing, or be infringing and need licensing. Why is audio different than images? Am I missing something here?

Law for music is different to other media types

Re: Stable-Audio-Demo

#216
post #213

Earlier quoted context omitted.

The reality is always a dynamic tension between law, regulation, precedent, and enforceability. It is possible to strangle OpenAI without strangling AI: pmarca is anti-OpenAI in print, but you can bet your butt he hopes to invest in whatever replaces it, and he’s got access to information that like, 10 people do. A useful example would be the Napster Wars: the music industry had been rent seeking (taking the fucking…

a16z are investors in openai

I'd look again: https://twitter.com/pmarca/status/1756803719327621141

Re: Stable-Audio-Demo

#217

Earlier quoted context omitted.

Tangential, but I tried to build chromium the other day but stopped when it said it required access to Google cloud platform to actually build it. If something requires a proprietary build system, does it matter that it's open source?

That is not true. See every distribution packaging chromium. In particular, this package[1] by openSUSE builds completely offline. Many other distributions require packages to build offline. [1] https://build.opensuse.org/package/show/network:chromium/chr...

I think I got my wires crossed with ChromiumOS which when I last read the docs seemed to suggest that Google cloud platform was required. I now can't find those specific docs either so I retract my statement.

Re: Stable-Audio-Demo

#218
post #205

As with Stable Diffusion, text prompting will be the least controllable way to get useful output with this model. I can easily imagine midi being used as an input with control net to essentially get a neural synthesizer.

I think it would be ideal if it could take the audio recording of humming or singing a melody together with a text prompt and spitting out a track that resembles it

1. Do your humming and pass it to something like Stable Audio with ControlNet

2. Convert/average the tone for each beat to generate something resembling a music sheet

3. Use vocaloid with LLM generated lyrics based on your prompt (or just put in your lyrics) and pass in the music file

4. Combine the 1-3

Would love to see this

Re: Stable-Audio-Demo

#219

Earlier quoted context omitted.

The majority of AI models out there (at least by popularity / capability) are proprietary; with weights and even model architectures that are treated as trade secret. Instead of having human-written music and movies that you legally can't copy, but practically can; you now have slop-generating models that live on a cloud server you have no control over. Artists and programmers who want to actually publish something -…

Is FSF's stance on AI actually clear? I thought they were just upset it was made by Microsoft. Creative Commons has been fairly pro-AI -- they have been quite balanced, actually, but they do say that opt-in is not acceptable, it should be opt-out at most. EFF is fairly pro AI too -- at least, against using copyright to legislate against it. You shouldn't discount progress in the open model ecosystem. You can sort of…

I haven't paid any attention to the FSF in years.

The Software Freedom Conservancy has been complaining about GitHub Copilot since 2022[0]. They specifically cite Copilot's use of training data in ways that violate the copyleft and attribution requirements of various FOSS licenses. Hector Martin (the guy porting Linux to MacBooks) also agrees with this. It's also important to note that the first AI training lawsuit was specifically to enforce GPL copyleft[1].

The EFF's argument has come across to me less like "AI is cool and good" and more like "copyright doesn't do a good job of protecting artists against AI taking their jobs". Cory Doctorow's also taken a similar position, arguing that unions are better at protecting against AI than copyright is. e.g. WGA being able to get contractual provisions preventing workers from being replaced with AI.

This is a different vein of opposition to AI from what we saw the following year in 2023 with artists and writers, though. Even then, those artists and writers aren't suddenly massively pro-copyright[2] and more consider it a means to fatally wound AI companies[3]. In contrast, big businesses that own shittons of copyright have been oddly quiet about AI. Sure, you have Getty Images and The New York Times suing Stability and OpenAI, but where's, say, Disney or Nintendo's litigation? These models can draw shittons of unlicensed fanart[4], and nobody cares. Wizards and Wacom made big statements against AI art, but then immediately got caught using it anyway, because stock image sites are absolutely flooded with it.

My personal opinion is that generative AI creates enough issues that we can't group them down into neat "pro-copyright" vs. "anti-copyright" arguments. People who share their work for free online are complaining about it while people who expect you to pay money for their work are oddly ambivalent. AI is orthogonal to copyright.

I will give you that the open model community is doing cool shit with their stolen loot. However, that's still something large corporations can benefit from (e.g. Facebook and LLaMA).

[0] https://sfconservancy.org/GiveUpGitHub/

[1] https://en.wikipedia.org/wiki/GitHub_Copilot#Licensing_contr...

[2] Which, for the record, many of them break.

[3] Their actual argument against AI is based on moral grounds, not legal ones. I don't think any artist is going to accept licensing payments for training data, they just want the models deleted off the Internet, full stop.

[4] OpenAI tried to ban asking for fanart, but if you ask for something vaguely related (e.g. "red videogame plumber" or "70s sci-fi robot") you'll get fanart every time.

Re: Stable-Audio-Demo

#220
post #203

Earlier quoted context omitted.

It isn't unusual for those in leadership positions to use such phrasing when talking about projects and products. It's not a "taking credit" from the engineers sort of thing, but rather about the leadership of the engineers.

Agreed. Leadership can sometimes bring actual value ;) And to be clear, I’m not sure Ed would call himself that. Those are my words, not his.

Ed here. Saw this thread and thought I'd weigh in.

Agreed, I wouldn't say I was hired to build Stable Audio. Crazy talented team of research engineers / software engineers / designers did the building.

Also wanted to clarify that I didn't quit due to concerns around the training data used for Stable Audio. I was proud of the approach we took to training data - a rev share with rights holders. I quit because of the prevailing view on training data at the wider company, as documented in its public response to the copyright office, where it argues that training on people's work without consent is fair use.

Post reply on HN