Live data from Hacker News

Stable Audio Open

stability.ai

81–90 of 137 posts

Re: Stable Audio Open

#81
post #11

> The new model was trained on audio data from FreeSound and the Free Music Archive. This allowed us to create an open audio model while respecting creator rights. This feels like the “Ethereum merge moment” for AI art. Now that there exists a prominent example with the big ethical obstacle (Proof of Work in the case of Ethereum, nonconsensual data-gathering in the case of generative AI) removed, we can actually have…

I also highly doubt anyone who signed agreements to have their music included in the Free Music Archive would have been OK with this. The particular type of license was important to contributors and there's a difference between allowing for rebroadcast without paying royalties and allowing for derivative works... I don't really care to argue the point, but it's why there were so many different types of licenses for t…

If you look at the repo where the model is actually hosted they specify

> All audio files are licensed under CC0, CC BY, or CC Sampling+.

These explicitly permit derivative works and commercial use.

> Attribution for all audio recordings used to train Stable Audio Open 1.0 can be found in this repository.

So it’s not being glossed over, and licenses are being abided by in good faith imo.

I wish they’d just added a sentence to their press release specifying this, though, since I agree it looks suspect if all you have to go by is that one line.

(Link: https://huggingface.co/stabilityai/stable-audio-open-1.0#dat... )

Re: Stable Audio Open

#82
post #55
post #36

Earlier quoted context omitted.

How do the rich become gradually richer under PoS? I'm flabbergasted by the level of math education. Assume we have 2 validators in the network; the first one owns 90% of the network, the second one owns 10%. Lets call them Whale and Shrimpy, respectively. To make the numbers round let's assume total circulating supply of ETH is 100 initially and that the yield resulting from being a validator is 10% per year. After…

Sure, friends also won't let friends skip the fact that circulating supply of ETH is now decreasing instead of increasing. Also, only ~30% tokens are staked. The 30% who chose to stake essentially tax the other 70% in use. Each of the validator do the same amount of work (ok, strictly speaking you get to do more when you have more ETH staked, but being a validator is cheap and does not cost significantly more energy…

> Also, only ~30% tokens are staked. The 30% who chose to stake essentially tax the other 70% in use.

And in PoW miners tax 100% of holders.

> what they receive is proportioned to how much they stake

Wealthy miners with state of the art ASICS benefit more than some kid mining at home with an old GPU. Maintenance/cost of mining equipment benefits from economies of scale too.

I hate being mean, but sorry, remembering to check one's assumption is a habit I gained after elementary school, so maybe that's too hard for you.

Re: Stable Audio Open

#83
post #11

> The new model was trained on audio data from FreeSound and the Free Music Archive. This allowed us to create an open audio model while respecting creator rights. This feels like the “Ethereum merge moment” for AI art. Now that there exists a prominent example with the big ethical obstacle (Proof of Work in the case of Ethereum, nonconsensual data-gathering in the case of generative AI) removed, we can actually have…

This has been Adobe Firefly's value proposition for months now. It works fine and is already being utilized in professional workflows with the blessing of lawyers.

Re: Stable Audio Open

#84

It produces decent audio, but something unpleasant about its high frequencies. And no voices, it doesn't seem to talk or sing. Udio, so far, is undefeated. And ElevenLabs' music demos were very very impressive, but it's still not released.

Have you tried suno? It is quite good at least for some genres

Suno is good at generating music, but its voices sound metallic with a dash of high-frequency noise. Which ruins it for me. It's almost there though, I think they will fix it in the next version.

Re: Stable Audio Open

#85
post #75
post #65

Earlier quoted context omitted.

Having exact copies of the samples inside the model weights would be an extremely inefficient use of space, and also it would not generalize, unless it generated a copy so close to the original that it would violate copyright law if used, I wouldn't find it very reasonable to think that there is a memorized copy inside the model weights somewhere.

A program that can produce copies is the same as a copy. How that copy comes into being (whether out of an algorithm or read from a support) is related, but not relevant.

>A program that can produce copies is the same as a copy.

A program that always produces copies is the same as a copy. A program that merely can produce copies categorically is not.

The Library of Babel[1] can produce copyrighted works, and for that matter so can any random number generator, but in almost every normal circumstance will not. The same is true for LLMs and diffusion models. While there are some circumstance that you can produce copies of a work, in natural use that's only for things that will come up thousands of times in its training set -- by and large, famous works in the public domain, or cultural touch-stones so iconic that they're essentially genericized (one main copyrighted example are the officially released promo materials for movies).

[1] https://libraryofbabel.info/

Re: Stable Audio Open

#86

It produces decent audio, but something unpleasant about its high frequencies. And no voices, it doesn't seem to talk or sing. Udio, so far, is undefeated. And ElevenLabs' music demos were very very impressive, but it's still not released.

None of these are impressive in the least. Anything I have heard from Udio is basically trash. It is the AI art equivalent of synthetic cats and pretty face shots. Who cares. What is ultimately going to be undefeated is training your own model.

I've heard many good things from Udio and the demo tracks from ElevenLabs are very high quality.

https://www.udio.com/songs/ai2uAaBffRGdWdTNNqAbDx

https://www.udio.com/songs/19xQAMG6E1UXG7wNvP7nDW

https://www.udio.com/songs/mPAFYyFgo7Nqjb8ypeFfh9

https://www.udio.com/songs/7sKM9jMwZrXwTTzmMYN9qv

https://www.udio.com/songs/coixNX1gnJ1oWT8z2LQddk

Re: Stable Audio Open

#87
post #57

Earlier quoted context omitted.

The idea that AI trained on artist created content is theft is kind of ridiculous anyway. Transformers aren't large archives of data with needles and thread to sew together pieces. The whole argument is meant to stifle an existential threat, not to halt some illegal transgression. If they cared about the latter a simple copyright filter on the output of the models would be all that's needed.

I fail to see how the argument is ridiculous; and I'll bet that a jury would find the idea that "there is a copy inside" at least reasonable, especially if you start with the premise that "the machine is not a human being." What you're left with is a machine that produces "things that strongly resemble the original, that would not have been produced, had you not fed the original into the machine." The fact that there…

If a made a bot that read amazon reviews and then output a meta-review for me, would that be a violation of amazon's copyright? (I'm sure somewhere in the amazon ToS they claim all ownership rights of reviews).

If it output those reviews verbatim, sure I can see the issue, the model is over fitting. But if I tweak the model or filter the output to avoid verbatim excerpts, does an amazon lawyer have a solid footing for a "violation of copyright" lawsuit?

Re: Stable Audio Open

#88
post #57

Earlier quoted context omitted.

The idea that AI trained on artist created content is theft is kind of ridiculous anyway. Transformers aren't large archives of data with needles and thread to sew together pieces. The whole argument is meant to stifle an existential threat, not to halt some illegal transgression. If they cared about the latter a simple copyright filter on the output of the models would be all that's needed.

I fail to see how the argument is ridiculous; and I'll bet that a jury would find the idea that "there is a copy inside" at least reasonable, especially if you start with the premise that "the machine is not a human being." What you're left with is a machine that produces "things that strongly resemble the original, that would not have been produced, had you not fed the original into the machine." The fact that there…

I am curious about models like encodec or soundstream. They are essentially meant to be codecs informed by the music they are meant to compress to achieve insane compression ratios. The decompression process is indeed generative since a part of the information that is meant to be decoded is in the decoder weights. Does that pass the smell test from a copyright law's perspective? I believe such a decoder model is powering gpt-4o's audio decoding.

Re: Stable Audio Open

#89
post #17

Earlier quoted context omitted.

Every time I call out the absurd interpretation of "Open Source" in this space in general, I get showered with downvotes and hateful attacks. One time someone posted their "AI mashup" project on reddit, that egregiously violated the terms of not only one, but several GPL-licensed projects. Calling this out earned me a lot of downvotes and replies with absolutely insane justifications from people with no clue. No one…

AI fans don't seem to care much for copyright, unless it's their work being stolen (remember the people that got mad at "prompt stealing"?). Companies are more risk-averse, though, and hobbyists on Reddit don't have the money to do anything serious with this software.

Nope. Most of the relevant software is OSS made by hobbyists. ComfyUI, llama.cpp, etc. For example, Nvidia is building stuff based on ComfyUI, a GPL-licensed application.

My complaint is about people ignoring the license of open source software made by hobbyists. I disagree with your ignorant "AI fans" generalization.

Re: Stable Audio Open

#90

Free idea because I’m never going to get around to building it: An “AI 8 track” app; click record and hum a melody, then add a prompt and click generate. The model converts your input to an instrument matching your prompt, keeping the same notes/rhythm you hummed in the original. Record up to 8 tracks and do some simple mixing. Would be a truly amazing thing for sketching songs! All you need is decent humming/singing…

Facebook’s MusicGen can pretty much do this
Post reply on HN