Live data from Hacker News

Stable Audio Open

stability.ai

91–100 of 137 posts

Re: Stable Audio Open

#91
post #21

When they released Stable Audio 2.0, I tried to create "unusual" songs with prompts like "roaring dragon tumbling rocks stormy morning". The results are quite interesting: https://www.youtube.com/@MarekGibney/videos I find it fascinating that you can put all information needed to recreate a whole complex song into a string like rough stormy morning car rocks hammering drum solo roaring dragon downtempo audiosparx-v2-…

That's a great example of the fact that information about something, say a song, isn't entirely encoded only in the medium you use to transfer it - it's partially there, and partially in the device you're using to read it! An MP3 file is just gibberish without a program that can decode it.

In this case, the whole album could indeed fit into a single TCP/IP packet - because the bulk of information that make up those songs is contained in the model, which weights however many gigabytes it does. The packet carrying your album is meaningless until the recipient also procures the model.

(Tangent: this observation was my first mind. blown. experience when reading GEB over a decade ago.)

Re: Stable Audio Open

#92
post #65

Earlier quoted context omitted.

Having exact copies of the samples inside the model weights would be an extremely inefficient use of space, and also it would not generalize, unless it generated a copy so close to the original that it would violate copyright law if used, I wouldn't find it very reasonable to think that there is a memorized copy inside the model weights somewhere.

An MP3 file is a lossy copy, but is still copyright infringement. Copyright infringement doesn't require exact copies.

I didn't say it takes an exact copy for copyright infringement.

Re: Stable Audio Open

#93
post #75
post #65

Earlier quoted context omitted.

Having exact copies of the samples inside the model weights would be an extremely inefficient use of space, and also it would not generalize, unless it generated a copy so close to the original that it would violate copyright law if used, I wouldn't find it very reasonable to think that there is a memorized copy inside the model weights somewhere.

A program that can produce copies is the same as a copy. How that copy comes into being (whether out of an algorithm or read from a support) is related, but not relevant.

Yeah that's right, I doubt that a model would generate an image or text so close to a real one to violate copyright law just by pure chance, the image/text space is incredibly large.

Re: Stable Audio Open

#94

Free idea because I’m never going to get around to building it: An “AI 8 track” app; click record and hum a melody, then add a prompt and click generate. The model converts your input to an instrument matching your prompt, keeping the same notes/rhythm you hummed in the original. Record up to 8 tracks and do some simple mixing. Would be a truly amazing thing for sketching songs! All you need is decent humming/singing…

App name - Beat-it https://www.youtube.com/watch?v=eZeYw1bm53Y https://www.nme.com/blogs/nme-blogs/the-incredible-way-micha...

A killer Ad would be converting each vocal track back into the original song

Re: Stable Audio Open

#95
post #11

> The new model was trained on audio data from FreeSound and the Free Music Archive. This allowed us to create an open audio model while respecting creator rights. This feels like the “Ethereum merge moment” for AI art. Now that there exists a prominent example with the big ethical obstacle (Proof of Work in the case of Ethereum, nonconsensual data-gathering in the case of generative AI) removed, we can actually have…

If you're worried about Proof of Work leading to giant server farms using huge amounts of energy, then I've got something to tell you about AI...

Re: Stable Audio Open

#96

Earlier quoted context omitted.

..which makes no sense. It is either an argument of ignorance or of purposeful deceit. There is no coherent data corpus (compressed or not) in ChatGPT. What is stored are weights that create a string of tokens that can recreate excerpts data that it was trained on, with some imperfect level of accuracy. Which I agree is problematic, and OpenAI doesn't have the right to disseminate that. But that doesn't mean OpenAI d…

> Content creators are doing a purposeful slight of hand to confabulate "outputting copyrighted data" with "training on copyrighted data". I don't think so, I think it's usually argued as two different things. The "training on copyrighted data" argument is usually that we never licensed this work for this sort of use and it is different enough from previously licensed uses that it should be treated differently. The "…

> Another argument is that licensed data is whitewashed by being run through a model. So you could have GPL licensed code that is open source run through a model and then output exactly the same but because it has been outputted by the model it is considered "cleaned" from the GPL restrictions. Clearly this output should still be GPL:ed.

I don't think anybody is making that argument. The NY Times claims to have gotten ChatGPT to spit out NY Times articles verbatim but there is considerable doubt about that. Regardless, everyone agrees that a verbatim (or close to) copy is copyright violation, even OpenAI. Every serious model has taken steps to prevent that sort of thing.

Re: Stable Audio Open

#97
post #11

> The new model was trained on audio data from FreeSound and the Free Music Archive. This allowed us to create an open audio model while respecting creator rights. This feels like the “Ethereum merge moment” for AI art. Now that there exists a prominent example with the big ethical obstacle (Proof of Work in the case of Ethereum, nonconsensual data-gathering in the case of generative AI) removed, we can actually have…

The "Etherium merge moment" is entirely different and it irks me to see it compared favorably with this project. It didn't 'solve' proof of work. It assigned positive value to past environmental harms. It prevents more proof of work, but doesn't solve past harms.

The only 'solution' (more a mitigation) to Etherium proof of work's environmental harms is to devalue it.

Unlike your example, this project actually seems to be a net positive for society that wasn't built on top of unnecessary harms.

Re: Stable Audio Open

#98
post #48

Earlier quoted context omitted.

..which makes no sense. It is either an argument of ignorance or of purposeful deceit. There is no coherent data corpus (compressed or not) in ChatGPT. What is stored are weights that create a string of tokens that can recreate excerpts data that it was trained on, with some imperfect level of accuracy. Which I agree is problematic, and OpenAI doesn't have the right to disseminate that. But that doesn't mean OpenAI d…

> There is no coherent data corpus (compressed or not) in ChatGPT. I disagree. If you can get the model to output an article verbatim, then that article is stored in that model. Just because it’s not stored in the same format is meaningless. It’s the same content regardless of whether it’s stored as plaintext, compressed text, PDF, png, or weights in a model. Just because you need an algorithm such as a specialized p…

>Just because you need an algorithm such as a specialized prompt to retrieve this memorized data, is also irrelevant.

I disagree. Granted I'm a layman and not a lawyer so I have no clue how the court feels. But I can certainly make very specialized algorithms to produce whatever output I want from whatever input I want, and that shouldn't let me declare any input as infringing on any rights.

For the reducto ad absurdum example: I demand everyone stops using spaces, using the algorithm 'remove a space and add my copyrighted text' it produces an identical copy of my copyrighted text.

For the less absurd example.. if I took any clean model without your copyrighted text, and brute forced prompts and settings until I produced your text, is your model violating the copyright or is my inputs?

Re: Stable Audio Open

#99
post #25
post #14

Earlier quoted context omitted.

Why is Proof of Work less ethical than Proof of Rich a.k.a. rich being gradually more rich without doing anything? Not saying PoW is safer (it's not), but less ethical is pretty a bold claim.

The environmental impact. And yes, I know, 0.5%, but my issue was always that if PoW currencies went from being a niche subculture to a point where it was used for everyday exchange (many people were arguing that this would and should happen) that 0.5% would surely go up by a great deal. To a point where crypto had to clear a super high bar of usefulness to counterbalance the harm it would do. To be fair, AI training…

This is greenwashing. You're still positively valuing the past harms from proof of work.

Solving the past harms means starting from scratch, not giving value to people who previously wasted electricity.

Re: Stable Audio Open

#100
post #57

Earlier quoted context omitted.

I fail to see how the argument is ridiculous; and I'll bet that a jury would find the idea that "there is a copy inside" at least reasonable, especially if you start with the premise that "the machine is not a human being." What you're left with is a machine that produces "things that strongly resemble the original, that would not have been produced, had you not fed the original into the machine." The fact that there…

If a made a bot that read amazon reviews and then output a meta-review for me, would that be a violation of amazon's copyright? (I'm sure somewhere in the amazon ToS they claim all ownership rights of reviews). If it output those reviews verbatim, sure I can see the issue, the model is over fitting. But if I tweak the model or filter the output to avoid verbatim excerpts, does an amazon lawyer have a solid footing fo…

As far as I understand, according to current copyright practices: If you sing a song that someone else has written, or pieces thereof, you are in violation. This is also the case of you switch out the instrumentation completely, say play trumpet instead of guitar, or a male choir sings a female line. If on would make a medley many such parts, it is not automatically not violation anymore either. So we do have examples of things being very far from verbatim copy, being considered violations.
Post reply on HN