Live data from Hacker News

Stable Audio Open

stability.ai

31–40 of 137 posts

Re: Stable Audio Open

#31
post #11

> The new model was trained on audio data from FreeSound and the Free Music Archive. This allowed us to create an open audio model while respecting creator rights. This feels like the “Ethereum merge moment” for AI art. Now that there exists a prominent example with the big ethical obstacle (Proof of Work in the case of Ethereum, nonconsensual data-gathering in the case of generative AI) removed, we can actually have…

The idea that AI trained on artist created content is theft is kind of ridiculous anyway. Transformers aren't large archives of data with needles and thread to sew together pieces. The whole argument is meant to stifle an existential threat, not to halt some illegal transgression. If they cared about the latter a simple copyright filter on the output of the models would be all that's needed.

[deleted]

Re: Stable Audio Open

#32

It produces decent audio, but something unpleasant about its high frequencies. And no voices, it doesn't seem to talk or sing. Udio, so far, is undefeated. And ElevenLabs' music demos were very very impressive, but it's still not released.

Have you tried suno? It is quite good at least for some genres

Re: Stable Audio Open

#33
post #11

> The new model was trained on audio data from FreeSound and the Free Music Archive. This allowed us to create an open audio model while respecting creator rights. This feels like the “Ethereum merge moment” for AI art. Now that there exists a prominent example with the big ethical obstacle (Proof of Work in the case of Ethereum, nonconsensual data-gathering in the case of generative AI) removed, we can actually have…

The idea that AI trained on artist created content is theft is kind of ridiculous anyway. Transformers aren't large archives of data with needles and thread to sew together pieces. The whole argument is meant to stifle an existential threat, not to halt some illegal transgression. If they cared about the latter a simple copyright filter on the output of the models would be all that's needed.

Our copyright model isn't sufficient yet. Is putting a work through a training/model sufficient to clear the transformative use bar? That doesn't make you safe from Trademarks. If the model can produce outputs on the other side that aren't sufficiently transformative then that single instance is a copyright violation.

Honestly, instead of trying to cleanup the output, it's much safer to create a licensed input corpus. People haven't because it's expensive and time consuming. Every time I engage with an AI vendor, my first question is do you indemnify from copyright violations of your output. I was shocked that Google Gemini/Bard only added that this year.

Re: Stable Audio Open

#34
post #28
post #23

Earlier quoted context omitted.

Proof of work mining is probably the only industry on the planet that has the potential to be carbon negative AND profitable.

(sigh) Is this the methane flaring argument, or Peter Thiel’s “windmills in Vermont”?

(sigh)

Do you expect a reply when you start like this?

Re: Stable Audio Open

#36
post #14
post #11

> The new model was trained on audio data from FreeSound and the Free Music Archive. This allowed us to create an open audio model while respecting creator rights. This feels like the “Ethereum merge moment” for AI art. Now that there exists a prominent example with the big ethical obstacle (Proof of Work in the case of Ethereum, nonconsensual data-gathering in the case of generative AI) removed, we can actually have…

Why is Proof of Work less ethical than Proof of Rich a.k.a. rich being gradually more rich without doing anything? Not saying PoW is safer (it's not), but less ethical is pretty a bold claim.

How do the rich become gradually richer under PoS? I'm flabbergasted by the level of math education.

Assume we have 2 validators in the network; the first one owns 90% of the network, the second one owns 10%. Lets call them Whale and Shrimpy, respectively.

To make the numbers round let's assume total circulating supply of ETH is 100 initially and that the yield resulting from being a validator is 10% per year. After the first year, 10 new ETH will have been minted. Whale would have gotten 9 ETH, and Shrimpy would have gotten 1 ETH. OP is assuming that as 9 is bigger than 1, Whale is getting richer faster than Shrimpy. But, let's look at the final situation globally.

At year 0:

Total ETH circulating supply: 100 ETH

Whale has 90 ETH. Owns 90% of the network.

Shrimpy has 10 ETH. Owns 10% of the network.

At year 1:

Total ETH circulating supply: 110 ETH

Whale has 99 ETH. Owns 90% of the network.

Shrimpy has 11 ETH. Owns 10% of the network.

Whale has exactly the same network ownership after validating for 1 whole year, the network is not centralizing at all! The rich are not getting richer any faster than the poor.

TL;DR: Friends don't let friends skip elementary math classes.

Re: Stable Audio Open

#37

Earlier quoted context omitted.

The idea that AI trained on artist created content is theft is kind of ridiculous anyway. Transformers aren't large archives of data with needles and thread to sew together pieces. The whole argument is meant to stifle an existential threat, not to halt some illegal transgression. If they cared about the latter a simple copyright filter on the output of the models would be all that's needed.

I think you should read the case material for NY Times v OpenAI and Microsoft. It literally says that within ChatGPT is stored, verbatim, large archives of NY Times articles and that they were able to retrieve them through their API.

..which makes no sense. It is either an argument of ignorance or of purposeful deceit. There is no coherent data corpus (compressed or not) in ChatGPT. What is stored are weights that create a string of tokens that can recreate excerpts data that it was trained on, with some imperfect level of accuracy.

Which I agree is problematic, and OpenAI doesn't have the right to disseminate that.

But that doesn't mean OpenAI doesn't have the right to train on it.

Content creators are doing a purposeful slight of hand to confabulate "outputting copyrighted data" with "training on copyrighted data".

It's illegal for me to read an NYT article and recite it from memory onto my blog.

It's not illegal for me to read an NYT article and write my own summary of the article's contents on my blog. This has been true forever and has forever been a staple in new content creation.

Re: Stable Audio Open

#38
post #11

> The new model was trained on audio data from FreeSound and the Free Music Archive. This allowed us to create an open audio model while respecting creator rights. This feels like the “Ethereum merge moment” for AI art. Now that there exists a prominent example with the big ethical obstacle (Proof of Work in the case of Ethereum, nonconsensual data-gathering in the case of generative AI) removed, we can actually have…

The idea that AI trained on artist created content is theft is kind of ridiculous anyway. Transformers aren't large archives of data with needles and thread to sew together pieces. The whole argument is meant to stifle an existential threat, not to halt some illegal transgression. If they cared about the latter a simple copyright filter on the output of the models would be all that's needed.

NY Times v OpenAI and Microsoft says the opposite, that verbatim, large archives of NY Times articles were retrieved via API. This may or may not matter to how LLMs work, but "large archive" seems accurate, other than semantic arguments (e.g. "Compressed archive" may be semantically more accurate).

Re: Stable Audio Open

#39
post #33

Earlier quoted context omitted.

The idea that AI trained on artist created content is theft is kind of ridiculous anyway. Transformers aren't large archives of data with needles and thread to sew together pieces. The whole argument is meant to stifle an existential threat, not to halt some illegal transgression. If they cared about the latter a simple copyright filter on the output of the models would be all that's needed.

Our copyright model isn't sufficient yet. Is putting a work through a training/model sufficient to clear the transformative use bar? That doesn't make you safe from Trademarks. If the model can produce outputs on the other side that aren't sufficiently transformative then that single instance is a copyright violation. Honestly, instead of trying to cleanup the output, it's much safer to create a licensed input corpus…

I'm honestly surprised AI-washing hasn't become way more widespread then it is at this point.

I mean recording a good song is hard. Generating a good song almost impossible. But my gut feeling would've been that recreating a popular song for plausible deniability would be a lot easier.

Same with republishing bestselling books and related media. (I.e. take Lord of the rings and feed it paragraph for paragraph into an LLM that you've prompted to rephrase each to a currently bestselling author.)

Re: Stable Audio Open

#40
> Warm arpeggios on an analog synthesizer with a gradually rising filter cutoff and a reverb tail

I appreciate that the underlying tech is completely different and much more powerful, but it is a pretty strange feeling to find a major AI lab's example sounding so similar to an actual Markov chain MIDI generator I made 14-15 years ago: https://youtu.be/depj8C21YHg?si=74a4DHP14EFCeYrB

(Not that similar, just enough for me to go "huh, what a coincidence").

Post reply on HN