Live data from Hacker News

Stable Diffusion Public Release

stability.ai

401–410 of 437 posts

Re: Stable Diffusion Public Release

#401
post #398

Earlier quoted context omitted.

An AI would need to generate: * Coherent video * Characters with backstory * Dialogue (including jokes and witty banter) * Music ...among many other things. Plus, the training set for video is orders of magnitude smaller than for digital art. (And is additionally burdened with copyright issues.) As I see it, there's simply no path from the DALL-E of today to something like that. And all for art that, essentially, "sa…

That's true. But the thing with technology, and the reason we've kept up with Moore's law is that someone eventually has a bright idea that leaves current methods and improvement extrapolation in the dust, and then the real thing happens earlier than the most optimistic dates, and performs better than what people expected. The question is not if one day an AI can generate a movie that you can't differentiate from a h…

>then the real thing happens earlier than the most optimistic dates

Is this like with self-driving cars that were supposed to be a consumer product in 5 years...in 2012?

Or like when Hinton said radiologists will be completely replaced in 5 years...in 2016?

By the way, Moore's law hasn't been a thing for a while in its original spirit.

Re: Stable Diffusion Public Release

#402
post #273

Earlier quoted context omitted.

What's wrong with Dell? Is Lenovo a bad option as well? Does NZXT offer on-site service?

You aren't getting on-site service with a consumer product. There are plenty of 3rd party people who can service your computer, though. It's like Lego. Dell uses crappy proprietary tech, poor quality components, and they have an all around bad reputation. NZXT uses good components and they make some of the best cases you can buy. I don't know much about Lenovo's desktop products. You might try posting on reddit.com/r…

> You aren’t getting on-site service with a consumer product.

Both Dell and Lenovo do this.

My daughter bought an Alienware laptop for college and when the keyboard broke, they sent a technician to her dorm to fix it (and she goes to school outside of the US).

If on-site service isn’t an option, what about a Mac Pro? At least with Apple I can take the machine to a store if I need to.

Re: Stable Diffusion Public Release

#403
post #98

Is there any way to download this on my PC and run it offline? Something like a command-line tool like $ ./something "cow flying in space" > cow-in-space.png that runs with local-only data (i.e. no internet access, no DRM, no weird API keys, etc like pretty much every AI-related application i've seen recently) would be neat.

Yes, that's actually the biggest reason this is such a cool announcement! You just need to download the model checkpoints from HuggingFace[0] and follow the instructions on their Github repo[1] and and you should be good to go. You basically just need to clone the repo, set up a conda environment, and make the weights available to the scripts they provide. [0] https://huggingface.co/CompVis/stable-diffusion [1] https…

What's the difference between those 4 checkpoints?

From the GitHub's README:

    sd-v1-1.ckpt: 237k steps at resolution 256x256 on laion2B-en. 194k steps at resolution 512x512 on laion-high-resolution (170M examples from LAION-5B with resolution >= 1024x1024).

    sd-v1-2.ckpt: Resumed from sd-v1-1.ckpt. 515k steps at resolution 512x512 on laion-aesthetics v2 5+ (a subset of laion2B-en with estimated aesthetics score > 5.0, and additionally filtered to images with an original size >= 512x512, and an estimated watermark probability 
Which one is the general use case checkpoint one should be using?

Re: Stable Diffusion Public Release

#404
post #33

The most interesting part, to me, of a release like this is the amount of "please don't abuse this technology" pleading. No licence will ever stop people from doing things that the licence says they can't. There will always be someone who digs into the internals and makes a version that does not respect your hopes and dreams. It's going to be bad. As I see it, within a couple years this tech will be so widespread and…

Isn't it good if people learn not to trust photos and images shared in social media and learn to treat them as entertainment. So much productivity gains :-)

Re: Stable Diffusion Public Release

#405

Earlier quoted context omitted.

Check out https://github.com/THUDM/CogVideo - progress is being made on coherent video generation. Characters and dialogue are effectively solved, just look at GPT-3. The entity behind StableDiffusion is also supporting generative music art, so let's see what is coming out of that: https://www.harmonai.org/ We are currently far away from generating a production quality movie with AI, but I don't think it's going to b…

>Characters and dialogue are effectively solved, just look at GPT-3. Is this the motherload of exaggeration? Current language models cannot generate coherent dialog (and even then it's mostly bad dialog) spanning more than a minute or two. And their current capabilities in that area are definitely significantly below those of the average human writer.

We were talking about a Marvel action flick, I don't think incredible dialog spanning multiple minutes is much of a thing apart from exposition dumps. I asked GPT-3 to spit out some paragraphs from a hypothetical script for Thor 5:

INT. DARKNESS We hear a faint beating heart. A moment later, we see a light slowly growing in the darkness. As the light grows, we see that it is coming from a glowing object in a person’s hand. The object is a hammer.

We see the face of the person holding the hammer. It is Thor. He looks tired and beaten.

Suddenly, we hear a voice from the darkness.

Black Panther: You are not welcome here, Thor.

Thor: I know. But I must speak with you.

Black Panther: You have nothing to say that I want to hear.

Thor: I come bearing a warning. Thanos is coming.

Black Panther: We are prepared.

Thor: He is not coming alone. He has an army.

Black Panther: So do we.

Thor: Thanos is not like any enemy you have faced before. He is ruthless and he will not stop until he has destroyed everything that you hold dear.

Black Panther: We will stop him.

Thor: I hope you can. Because if you cannot, then all is lost.

Eh, looks real enough to me. Fine tune the model with all the specialities that make up Marvel movies and you'll crank out good-enough drafts in no time.

Re: Stable Diffusion Public Release

#406

Earlier quoted context omitted.

>Characters and dialogue are effectively solved, just look at GPT-3. Is this the motherload of exaggeration? Current language models cannot generate coherent dialog (and even then it's mostly bad dialog) spanning more than a minute or two. And their current capabilities in that area are definitely significantly below those of the average human writer.

We were talking about a Marvel action flick, I don't think incredible dialog spanning multiple minutes is much of a thing apart from exposition dumps. I asked GPT-3 to spit out some paragraphs from a hypothetical script for Thor 5: INT. DARKNESS We hear a faint beating heart. A moment later, we see a light slowly growing in the darkness. As the light grows, we see that it is coming from a glowing object in a person’s…

>cannot generate coherent dialog (and even then it's mostly bad dialog) spanning more than a minute or two

I think that was pretty clear and that posted dialog is a perfect illustration.

You cannot generate the entire movie script coherently without significant human input and that's not going to change in the next several years. So, your initial claim that dialogue is "solved" is indeed false.

Re: Stable Diffusion Public Release

#407
post #67

"This release is the culmination of many hours of collective effort to create a single file that compresses the visual information of humanity into a few gigabytes." If something like this is possible, does this mean there's actually far less meaningful information out there than we think? Could you in fact pack virtually all meaningful information ever gathered by humanity onto a 1TiB or smaller hard drive? Obviousl…

You can pack virtually all meaningful information ever gathered by humanity onto a single bit, but its gonna be lossy. And what is your definition of "meaningful information" anyway. Meaningful today might not be meanigful yesterday. Nobody cares about spin of each electron in my brain today, but in 4 centuries my descendants will be like "if only we had that information, we could simulate our great-...-great parent…

Lossy compression isn’t linear.

If you play with JPEG quality you’ll see that the difference is barely perceptible for a while and then if you keep going down it becomes very noticeable.

So what does this look like with a general model?

Re: Stable Diffusion Public Release

#408
post #3

This release changes society forever. Free and open access to generate a hyper-realistic image via just a text prompt is more powerful than I think we can imagine currently. Art, media, politics, conspiracy theories; all of it changes with this.

What will change society forever is when the hardware required to run this software is available in the latest medium/high end phone and 100's million of people can download the app to say what they want.

Re: Stable Diffusion Public Release

#409
post #88

Moral question: Such tool will be used to generate lewd imagery involving virtual minors. No way to prevent it upstream by outlawing feeding it real content (whose possession already is illegal). Suffice to add 'childrenize' layer onto adult NN or something. How will the legal system react ? Bundle it into illegal imagery, period ? Maybe it's already the case - I think drawing made public is, not sure. If not, on wha…

Won’t somebody please think of the imaginary children?

Re: Stable Diffusion Public Release

#410
post #46

So is there a simple way to do this online? I have no dedicated gpu and won’t buy one just for this. We‘d pay.

https://beta.dreamstudio.ai/

10 pounds just evaporated into nothingness and a few images barely above dalle-mini...
Post reply on HN