Live data from Hacker News

Stable Audio Open

stability.ai

41–50 of 137 posts

Re: Stable Audio Open

#41
post #30
post #25

Earlier quoted context omitted.

The environmental impact. And yes, I know, 0.5%, but my issue was always that if PoW currencies went from being a niche subculture to a point where it was used for everyday exchange (many people were arguing that this would and should happen) that 0.5% would surely go up by a great deal. To a point where crypto had to clear a super high bar of usefulness to counterbalance the harm it would do. To be fair, AI training…

There is no "environmental impact". Environmental impact comes from energy production, not energy usage. It's incoherent to argue others should tamper down their energy usage because most folks producing energy aren't doing it in an ethical way.

Ultimately any proof-of-work system has to burn joules rather than clock cycles (because any race on cycles-per-joule is rapidly caught up), and that makes it clearer where the waste is: to be economically stable, in the face of adversarial actions by other nation states who sometimes have a vested interest in undermining your currency so actively seek the chaos and loss of trust in a double-spend event, your currency has to be backed by more electricity than any hostile power can spend on breaking it.

Re: Stable Audio Open

#42

From announcement I couldn’t figure out if it can do audio to audio. Text to audio is too limiting. I’d rather input a melody or a drum beat and have the AI compose around it.

This kind of exists, but I doubt there are any commercial solutions based on it yet. https://crfm.stanford.edu/2023/06/16/anticipatory-music-tran...

Their paper says that they trained it on the Lakh MIDI dataset, and they have a section on potential copyright issues as a result.

Assuming you don't care for legal issues, theoretically you could do: raw signal -> something like Spotify Basic Pitch (outputs MIDI) -> Anticipatory (outputs composition) -> Logic Pro/Ableton/etc + Native Instruments plugin suite for full song

Re: Stable Audio Open

#44
post #12

Free idea because I’m never going to get around to building it: An “AI 8 track” app; click record and hum a melody, then add a prompt and click generate. The model converts your input to an instrument matching your prompt, keeping the same notes/rhythm you hummed in the original. Record up to 8 tracks and do some simple mixing. Would be a truly amazing thing for sketching songs! All you need is decent humming/singing…

Google MusicLM (and probably lots of other tools) do this: "MusicLM .. can transform whistled and hummed melodies according to the style described in a text caption." https://google-research.github.io/seanet/musiclm/examples/

Sounds like Suno AI will soon have this feature as well https://x.com/suno_ai_/status/1794408506407428215

Re: Stable Audio Open

#45
post #34
post #28

Earlier quoted context omitted.

(sigh) Is this the methane flaring argument, or Peter Thiel’s “windmills in Vermont”?

(sigh) Do you expect a reply when you start like this?

I don’t really care either way. I’m tired of having to debunk the same sloppy arguments year after year.

Re: Stable Audio Open

#46

Earlier quoted context omitted.

I think you should read the case material for NY Times v OpenAI and Microsoft. It literally says that within ChatGPT is stored, verbatim, large archives of NY Times articles and that they were able to retrieve them through their API.

..which makes no sense. It is either an argument of ignorance or of purposeful deceit. There is no coherent data corpus (compressed or not) in ChatGPT. What is stored are weights that create a string of tokens that can recreate excerpts data that it was trained on, with some imperfect level of accuracy. Which I agree is problematic, and OpenAI doesn't have the right to disseminate that. But that doesn't mean OpenAI d…

When you describe ChatGPT as just a model with weights that can create a string of tokens, is it any different from any lossless compression algorithm?

I'm sure if I had a JPEG of some copyrighted raw image it could still be argued that it is the same image. JPEG is imperfect, the result you get is the same every time you open it but it's not the same as the original input data.

ChatGPT would give you the same output every time, and it does if you turn off the "temperature" setting. Introduce a bit of randomness into a JPEG decoder and functionally what's the difference? A slightly different string of tokens for ChatGPT versus a slightly different collection of pixels for a JPEG.

Re: Stable Audio Open

#47
post #30
post #25

Earlier quoted context omitted.

The environmental impact. And yes, I know, 0.5%, but my issue was always that if PoW currencies went from being a niche subculture to a point where it was used for everyday exchange (many people were arguing that this would and should happen) that 0.5% would surely go up by a great deal. To a point where crypto had to clear a super high bar of usefulness to counterbalance the harm it would do. To be fair, AI training…

There is no "environmental impact". Environmental impact comes from energy production, not energy usage. It's incoherent to argue others should tamper down their energy usage because most folks producing energy aren't doing it in an ethical way.

It seems you've come up with a proof that there's no such thing as wasting electricity. When you prove an extraordinary claim like that, it's time to go back and figure out how you got it wrong.

Re: Stable Audio Open

#48

Earlier quoted context omitted.

I think you should read the case material for NY Times v OpenAI and Microsoft. It literally says that within ChatGPT is stored, verbatim, large archives of NY Times articles and that they were able to retrieve them through their API.

..which makes no sense. It is either an argument of ignorance or of purposeful deceit. There is no coherent data corpus (compressed or not) in ChatGPT. What is stored are weights that create a string of tokens that can recreate excerpts data that it was trained on, with some imperfect level of accuracy. Which I agree is problematic, and OpenAI doesn't have the right to disseminate that. But that doesn't mean OpenAI d…

> There is no coherent data corpus (compressed or not) in ChatGPT.

I disagree.

If you can get the model to output an article verbatim, then that article is stored in that model.

Just because it’s not stored in the same format is meaningless. It’s the same content regardless of whether it’s stored as plaintext, compressed text, PDF, png, or weights in a model.

Just because you need an algorithm such as a specialized prompt to retrieve this memorized data, is also irrelevant. Text files need to be interpreted in order to display them meaningfully, as well.

Re: Stable Audio Open

#49
post #33

Earlier quoted context omitted.

The idea that AI trained on artist created content is theft is kind of ridiculous anyway. Transformers aren't large archives of data with needles and thread to sew together pieces. The whole argument is meant to stifle an existential threat, not to halt some illegal transgression. If they cared about the latter a simple copyright filter on the output of the models would be all that's needed.

Our copyright model isn't sufficient yet. Is putting a work through a training/model sufficient to clear the transformative use bar? That doesn't make you safe from Trademarks. If the model can produce outputs on the other side that aren't sufficiently transformative then that single instance is a copyright violation. Honestly, instead of trying to cleanup the output, it's much safer to create a licensed input corpus…

Nothing will ever protect you from trademark violations because trademarks can be violated completely by accident without knowledge of the real work. Copying is not the issue.

Re: Stable Audio Open

#50

Earlier quoted context omitted.

..which makes no sense. It is either an argument of ignorance or of purposeful deceit. There is no coherent data corpus (compressed or not) in ChatGPT. What is stored are weights that create a string of tokens that can recreate excerpts data that it was trained on, with some imperfect level of accuracy. Which I agree is problematic, and OpenAI doesn't have the right to disseminate that. But that doesn't mean OpenAI d…

When you describe ChatGPT as just a model with weights that can create a string of tokens, is it any different from any lossless compression algorithm? I'm sure if I had a JPEG of some copyrighted raw image it could still be argued that it is the same image. JPEG is imperfect, the result you get is the same every time you open it but it's not the same as the original input data. ChatGPT would give you the same output…

Did you mean lossy compression algorithm? That would make sense.
Post reply on HN