Live data from Hacker News

Stable Diffusion 2.0

stability.ai

481–490 of 519 posts

Re: Stable Diffusion 2.0

#481

It kind of annoys me that they removed NSFW images from the training set. Not because I want to generate porn (though some people do), but because I feel that they're foisting a puritan ethic on me. I don't consider the naked body inherently bad, and I don't like seeing new technology carry this (wrong, in my opinion) stigma. Then again, it's their model, they can do whatever they want with it, but it still leaves me…

It’s annoying when you get NSFW results when you didn’t ask for them, so it may be better to segregate them.

But they aren't segregating them. They didn't release two models, one SFW and one NSFW. They segregated them before, with the filter you could disable, but now it's all SFW-only.

Re: Stable Diffusion 2.0

#482
post #449

Earlier quoted context omitted.

Human artists derive their inspiration and styles from a large set of copyrighted works, but they are free to produce new art despite of that. Art would have developed much slower and be much poorer if, for example, Impressionism or Cubism had been entangled in long ownership confrontations in courts. Then there's the fact that humanity has been able to develop and share art and literary works for thousands of years…

Human artists cannot produce thousands of works in a few hours. This arguments come up in every thread, and I'm baffled that people don't think the scale matters. You may also be observed in public areas by police, but it would be an orwellian dystopia to have millions of cameras in spaces analyzing everyone's behavior in public. Scale matters. (But I'm indeed in favor of weaker copyright laws! But preferably to take…

> it would be an orwellian dystopia to have millions of cameras in spaces analyzing everyone's behavior in public.

Aren't there already 80M+ surveillance cameras in the US?

Outside of the US, London seems to have a lot of CCTV cameras.

Do privacy laws restrict how they can be used and whether they can be monitored by AI systems?

Re: Stable Diffusion 2.0

#483
post #204

The crimes against the creative people are getting better and better. What a time to live, when your entire career burns to dust just because. I hope AI gets these programmers jobs soon. Then we all can go to the woods and have a good life, finally.

Thanks for your premature concern, but we'll be fine. Despite how it may appear to a layperson such as yourself, the value of human creativity is in no way diminished by the release of this tool or others like it.

Ok. The irony. Actually, after 20+ years in the tech industry, I will say this:

Your beloved corporations don't have a metric called “creativity”, they have a bottom line, and she has all the powers.

I am an artist by education and can confirm that creativity is overrated, the processes that artist follows and repetition towards a given goal deliver the results.

Whatever feelings or ideas you have, the actual craft is the medium in which you will deliver.

Reducing *The Path* to text input is not an artistic or craftsmanship process.

There is no creativity involved. May be, someone with more knowledge about the real process and broader visual culture will make more aesthetically right choices. But this can be automated too.

This is not a “tool”, like Photoshop. This is something else. And all of you know this.

More than 50 percent of frontend code is boilerplate. CRUD apps follow similar logic. Why not automate this repetitive processes first?

No. Corporations are starting the automation from the lowest risk crowd—the digital artists, they have low representation, no coherent community and are always ready to sell themselves for pennies.

Now they will compete with the machines. And your time in this battle will come. Soon.

Re: Stable Diffusion 2.0

#484

Earlier quoted context omitted.

It’s annoying when you get NSFW results when you didn’t ask for them, so it may be better to segregate them.

But they aren't segregating them. They didn't release two models, one SFW and one NSFW. They segregated them before, with the filter you could disable, but now it's all SFW-only.

Eh, fine-tuning seems to work well enough that it can be added back in after.

Though, previous fine-tunings/textual inversions won’t work since the CLIP encoder has been replaced too. I’d be interested in knowing if it needs to be retrained too for this case.

Re: Stable Diffusion 2.0

#485
post #166

Earlier quoted context omitted.

Did they exclude celebrities, politicians, and religious and political symbols? Deceitful extremists and vengeful criminals fabricating lies seem to be a far more serious problem than fantasy porno.

That's a really interesting point, and it makes me realize that the Nancy Reagan 'what constitutes porn' question is obviously super old and problematic. Also lexica.art is swarming with celebrity fantasy porn that just has a thin stylistic filter of paintings from the 19th century. And a plethora of furry daddies that you can't not love. I get why these models should be curated but I also like that the sketchy porn…

> Nancy Reagan 'what constitutes porn' question

I thought that was Justice Stewart? And then he answered it "I know it when I see it."

Re: Stable Diffusion 2.0

#486

Wow, just wow! Newbie question, why can’t someone just take a pre-trained model/network with all the settings/weights/whatever and run it on a different configuration (at a heavily reduced speed)? Isn’t it like a Blender/3D studio/Autocad file, where you can take the original 3D model and then render it using your own hardware? With my single GOU it will take days to raytrace a big scene, whereas someone with multipl…

If you use the provided pytorch code, have a modern CPU and enough physical RAM, you can do this currently. As you suggest, inference/generation will take anywhere from hours to days using a CPU instead of a GPU or other ML-accelerator-chip.

Re: Stable Diffusion 2.0

#487

Earlier quoted context omitted.

The code is open source, the model is a data file that the open source code operates on. It's similar to engine recreations for old games (OpenRCT, OpenTTD) that use original, proprietary assets to play the games with their open source engines. Similar to those games, anyone is also able to distribute their own open data files if they so wish It's unlikely anyone actually will start training an open source AI model f…

I don't know what your point is. They use the terms "open source AI models" and "open source Generative AI models" Yes, someone else could spend the millions of dollars to create a model that actually is open source, but shouldn't the people advertising their models as open source do that?

Should they distribute the data files according to the open source standards? Maybe. "Open" does not mean "open source", though; "open data" does not necessarily allow unlimited access and use of such data available, it's usually behind some kind of ToS document nobody reads and an API key. Applying open source expectations to anything with open in the name will often leave you disappointed outside the FOSS world.

Does not openly distributing their data files make their code any less open source? I don't think so. The code is open and licensed with a FOSS license. They spend time and money on creating a model and give the world the ability to replicate their model if it can collect the necessary funds. There are plenty of other open source projects that require vast arrays of server racks and compute power to be useful, that doesn't change anything about the openness of the code.

Re: Stable Diffusion 2.0

#488
post #250

Earlier quoted context omitted.

It confused me that the letter boxes were divided in 7+3, thus I thought it would be two words while the correct answer was a single 10 letter word. Maybe try to avoid wrapping words.

Nice Observation!! I'm thinking including a start and end mark to improve the UX would work well. I can't avoid wrapping as the prompt might be very large.

Put a hyphen at the end!

Re: Stable Diffusion 2.0

#489
post #358

Earlier quoted context omitted.

I'm disappointed they didn't push parameter count higher, but I suppose they want to maintain the ability to run on older/lower end consumer GPUs. Unfortunately it severely limits how high-quality the output can be.

They're motivating that choice via this paper: https://arxiv.org/pdf/2203.15556.pdf The paper shows that you can get better performance than gpt-3 with a much smaller model if you bump up the training time and training data like x4.

Larger models are still much better. Google's parti model can do text perfectly and follows prompts way more accurately than Stable Diffusion. It's 20B parameters and with the latest int8 optimizations it should be possible to get that running on a consumer 24GB card in theory.

I think they're looking into larger models later though

Re: Stable Diffusion 2.0

#490
post #240
post #176

Earlier quoted context omitted.

So does Pixelmator. You can try the free trial which comes with this feature.

Pixelmator's Super ML Resolution does a great job with upscaling images, can highly recommend it.

Yeah, I use it quite a lot on stuff generated by Midjourney and the results are always great.
Post reply on HN