Live data from Hacker News

Stable-Audio-Demo

stability-ai.github.io

141–150 of 249 posts

Re: Stable-Audio-Demo

#141
The few examples I was able to play are very promising, unfortunately the host seems to be getting some sort of HN-hug, because all the audio files are buffering every other second -- they seem to throttle at 32 KiB/s.

Re: Stable-Audio-Demo

#142
post #128

Earlier quoted context omitted.

That makes no sense. OpenAI must lose and it must not be possible to have proprietary models based on copyrighted works. It's not fair use because OpenAI is profiting from the copyright holders work and substituting for it while not giving them recompense. The alternative is that any models widely trained on copyrighted work are uncopyrightable and must be disclosed, along with their data sources. In essence this is…

Just because something is not copyrightable doesn’t automatically mean it must be disclosed. If weights aren’t copyrightable (and I don’t think they should be, as the weights are not a human creation), commercial AI’s just get locked behind API barriers, with terms of usage that forbid cloning. Copyright then never enters the picture, unless weights get leaked. Whether or not that’s equitable is in the eye of the beh…

> Just because something is not copyrightable doesn’t automatically mean it must be disclosed.

No I'm saying that's what they law should be, because models can be built and used without anyone knowing. If it's illegal not to disclose them you can punish people.

Copyright is something that protects the little guy as much as big corps. But the former has more to lose as a group in the world of AI models, and they will lose something here no matter what happens.

Re: Stable-Audio-Demo

#143
Music without changes is boring. I enjoyed the much less stable results of OpenAI's JuleBox (2021?) more than any music AI to come since. Their sound quality is better but they only seem to produce one monotonous texture at a time.

Re: Stable-Audio-Demo

#144

Earlier quoted context omitted.

> If you require licensing fees for training data, you kill open source ML. And likely proprietary ML as well, hopefully. (To be clear, I think AI is an absolutely incredible innovation, capable of both good and harm; I also think it's not unreasonable to expect it to play a safer, slower strategy than the Uber "break the rules to grow fast until they catch up to you" playbook.) I'm all for eliminating copyright. Unt…

AI is a genie that you can't really stuff back into a bottle. It's out and it's global. If the US had tighter regulations, China or someone else will take over the market. If AI is genuinely transformative for productivity, then the US would just fall behind, sooner or later.

Then let them! If another country put forward tighter regulations to help actual people over and above the state that holds them, then that is good in itself, and either way will pay for itself. Why are we worried about China or whoever taking over the market of something that we see has bad effects?

Like, we see this line everywhere now, and it simply doesnt make sense. At some point you just have to believe something, be principled. Treating the entire world as this zero sum deadlock of "progress" does nothing but prevent one from actually being critical about anything.

This would-be Oppenheimer cosplay is growing really old in these discussions.

Re: Stable-Audio-Demo

#145

> Warning: This website may not function properly on Safari. For the best experience, please use Google Chrome. We've come full circle with the 90's and Internet Explorer. Well I guess this time the dominant browser is opensource so that's atleast something... Can someone please create an animated GIF button for Chrome which says: "Best viewed with Google Chrome"?

> Can someone please create an animated GIF button for Chrome which says: "Best viewed with Google Chrome"?

Here you go:

Edit: View the button: https://indiscipline.github.io/post/best-viewed-in-google-ch...

Re: Stable-Audio-Demo

#146
post #92

Earlier quoted context omitted.

I tried generating music on stableaudio.com and, yes, it's bad. However, given the blistering pace of developing in these models, I would not be surprised if these sound incredible in a year or two.

Everyone every time seems to assume a linear (or exponential) curve upwards. But what is the proof for that? I consider it far more likely that we had a breakthrough and now rushing towards the next plateau. Maybe are nearing that. Like in the curve of a PID controller. It's how most or many human improvements go.

The plateau we're heading for is getting professional human level output from these models with logarithmic progress.

I suspect this is because the underlying production factors like compute, data & model design are steadily improving whilst humans have diminishing sensitivity to output quality.

In the game of AI generated photorealistic images or history essays there's not much improvement left to make. Most humans are already convinced by the output of these things.

Re: Stable-Audio-Demo

#147

Earlier quoted context omitted.

For generative models, if the model authors do not publish the architecture of their model; and, the model uses a transformation from text to another kind of media; you can assume that they have delegated some part of their model to a text encoder or similar feature which is trained on data that they do not have an express license to. Even for rightsholders with tens of millions to hundreds of millions of library ite…

If you require licensing fees for training data, you kill open source ML. That’s why it’s important for OpenAI to win the upcoming court cases. If they lose, they’ll survive. But it will be the end of open model releases. To be clear, I don’t like the idea of companies profiting off of people’s work. I just like open source dying even less.

The reality is always a dynamic tension between law, regulation, precedent, and enforceability.

It is possible to strangle OpenAI without strangling AI: pmarca is anti-OpenAI in print, but you can bet your butt he hopes to invest in whatever replaces it, and he’s got access to information that like, 10 people do.

A useful example would be the Napster Wars: the music industry had been rent seeking (taking the fucking piss really) for decades and technology destroyed the free ride one way or another. The public (led by the technical/hacker/maker public) quickly showed that short of disconnecting the internet, we were going to listen to the 2 good songs without buying the 8 shitty ones. The technical public doesn’t flex its muscles in a unified way very often, but when it does, it dictates what is and isn’t on the menu.

The public wants AI, badly. They want it aligned by them within the constraints of the law (which is what “aligned” should mean to begin with).

The public is getting what it wants on this: you can bet the rent. Whether or not OpenAI gets on board or gets run the fuck over is up to them.

“You in the market for a Tower Records franchise Eduardo?”

Re: Stable-Audio-Demo

#148

Just a few days ago I was down voted for stating AI will be better in creating music than human would be: https://news.ycombinator.com/item?id=39273380#39273532 Now this is released and now I feel I got grist to my mill. Sure it still kind of sucks, but it's very impressive for a _demo_. Remember that this tech is very much in it's infancy and it's very impressive already.

I don't find this music to be good in any way. It sounds interesting over a few notes, but then completely fails to find any kind of progression that goes anywhere interesting, never iterating on the theme, never teasing you with subtle or surprising variation over a core theme, no built-ups or clear resolution. Very annoying to actually listen to.

Re: Stable-Audio-Demo

#149

Earlier quoted context omitted.

> If you require licensing fees for training data, you kill open source ML. And likely proprietary ML as well, hopefully. (To be clear, I think AI is an absolutely incredible innovation, capable of both good and harm; I also think it's not unreasonable to expect it to play a safer, slower strategy than the Uber "break the rules to grow fast until they catch up to you" playbook.) I'm all for eliminating copyright. Unt…

It’s not so clear cut. Many lawyers believe all that matters is whether the output of the model is infringing. As much as people love to cite ChatGPT spitting out code that violates copyright, the vast majority of the outputs do not. Those that do, are quickly clamped down on — you’ll find it hard to get Dalle to generate an image of anything Nintendo related, unless you’re using crafty language. There’s also the mor…

> Many lawyers believe all that matters is whether the output of the model is infringing.

What I don't understand (as a European with little knowledge of court decisions on fair use): with the same reasoning you might make software piracy a case of 'fair use', no? You take stuff someone else wrote - without their consent - and use it to create something new. The output (e.g. the artwork you create with Photoshop) is definitely not copyrighted by the manufacturer of the software. But in the case of software piracy, it is not about the output. With software, it seems clear that the act of taking something you do not have the rights for and using it for personal (financial) gain is not covered by fair use.

Why can OpenAI steal copyrighted content to create transformative works but I cannot steal Photoshop to create transformative works? What am I missing?

Re: Stable-Audio-Demo

#150
post #8

As with Stable Diffusion, text prompting will be the least controllable way to get useful output with this model. I can easily imagine midi being used as an input with control net to essentially get a neural synthesizer.

Yes. Since working on my AI melodies project ( https://www.melodies.ai/ ) two years ago, I've been saying that producing a high-quality, finalized song from text won't be feasible or even desirable for a while, and it's better to focus on using AI in various aspects of music making that support the artist's process.

Emad hinted here on HN the last time this was discussed that they were experimenting with exactly that. It will come, by them or by someone else quickly.

Text-prompting is just a very coarse tool to quickly get some base to stand on, ControlNet is where the human creativity again enters.

Post reply on HN