Live data from Hacker News

Stable Diffusion 2.0

stability.ai

381–390 of 519 posts

Re: Stable Diffusion 2.0

#381

Earlier quoted context omitted.

Interesting project but terrible naming.

To be clear to the original poster, the naming is terrible because of the nazi associations of the number 88, correct?

I didnt know that, 88 is associated with good luck, fortune, and money in chinese. so you see 88 everything in chinese.

Re: Stable Diffusion 2.0

#382

Earlier quoted context omitted.

The point is that there's is no practical limit on compression. You don't need "AI" or anything besides very basic statistics to get astronomical compression ratios. (See: "zip bomb".) The only practical limit is the amount of information entropy in the source material, and if you're going to claim that internet pictures are particularly information-dense I'd need some evidence, because I don't believe you.

Correct, however "compression is equivalent to general intelligence" ( http://prize.hutter1.net/hfaq.htm#compai ) and so in a sense, all learning is compression. In this case, SD applies a level of compression that is so high that the only way it can sustain information from its inputs is by capturing their underlying structure. This is a fundamentally deeper level of understanding than image codecs, which merely cap…

I fail to see the difference between "underlying structure" and "short-range visual features".

Both are just simple statistical relationships between parameters and random variables.

Re: Stable Diffusion 2.0

#383
post #163
post #131

Hopefully related: If I'm a photographer wanting to improve resolution of my content for printing, what's my current best bet for upscaling? Is it realistic to make use of this on the command line, feeding it my own images? Or has someone wrapped it in an app or online service?

As a counter-recommendation, Topaz’s much-advertised Gigapixel AI is rarely useful. Their Denoise and Sharpen apps are good though.

I spent a lot of time last month using Gigapixel (actually the improved version in their new Photo AI product) last month on dozens of images for my dad's memoir. There were a couple failures where the input image was just so blurry or low-res that it couldn't be saved, but Topaz significantly improved image quality while upscaling in 90+% of cases.

Re: Stable Diffusion 2.0

#384
post #163

Earlier quoted context omitted.

As a counter-recommendation, Topaz’s much-advertised Gigapixel AI is rarely useful. Their Denoise and Sharpen apps are good though.

I dunno - I've found it useful on a bunch of images[1] but I tend to try Pixelmator Pro first because that's a simple key combination to enlarge an image and 90% of the time it's Good Enough for my purposes. [1] The new Photo AI, on the other hand, is slow, clunky, and not infrequently glitches out wildly. But on the plus side it does combine sharpening and denoising into one workflow.

I was super unimpressed with the 1.0 release of Photo AI; in particular, the sharpening was a LOT slower than standalone. But that's fixed now, and unless Topaz starts backporting the improved models to the standalone tools -- so far, they have not -- Photo AI will get you better results.

Re: Stable Diffusion 2.0

#385

Earlier quoted context omitted.

Correct, however "compression is equivalent to general intelligence" ( http://prize.hutter1.net/hfaq.htm#compai ) and so in a sense, all learning is compression. In this case, SD applies a level of compression that is so high that the only way it can sustain information from its inputs is by capturing their underlying structure. This is a fundamentally deeper level of understanding than image codecs, which merely cap…

I fail to see the difference between "underlying structure" and "short-range visual features". Both are just simple statistical relationships between parameters and random variables.

Sure, but why would that not apply to humans? And we don't consider it copyright violation if a human learns painting by looking at art.

Re: Stable Diffusion 2.0

#386

Is there a good explanation of how to train this from scratch with a custom dataset[0]? I've been looking around the documentation on Huggingface, but all I could find was either how to train unconditional U-Nets[1], or how to use the pretrained Stable Diffusion model to process image prompts (which I already know how to do). Writing a training loop for CLIP manually wound up with me banging against all sorts of stra…

Here’s a tutorial on how to fine tune stable diffusion form the guy who made text-to-pokemon:

https://lambdalabs.com/blog/how-to-fine-tune-stable-diffusio...

Re: Stable Diffusion 2.0

#387
post #363

In addition to removing NSFW images from the training set, this 2.0 release apparently also removed commercial artist styles and celebrities [1]. While it should be possible to fine tune this model to create them anyway using DreamBooth or a similar approach, they clearly went for the safe route after taking some heat. 1. https://twitter.com/emostaque/status/1595731407095140352?s=4...

Removing NSFW content is fine, people who care about that can work around it easily. Removing celebrities and commercial artists was a mistake though and I expect this will need to be really impressive in other ways or people aren't going to bother using it.

Re: Stable Diffusion 2.0

#388
post #363

In addition to removing NSFW images from the training set, this 2.0 release apparently also removed commercial artist styles and celebrities [1]. While it should be possible to fine tune this model to create them anyway using DreamBooth or a similar approach, they clearly went for the safe route after taking some heat. 1. https://twitter.com/emostaque/status/1595731407095140352?s=4...

Mixing artist names was by far the most effective way to create aesthetically pleasing images, this is a huge change. DreamBooth can only fine-tune on a couple dozen images, and you can't train multiple new concepts in one model, but maybe someone will do a regular fine-tune or train a new model.

Re: Stable Diffusion 2.0

#389
post #27

I am a solo dev working on a creative content creation app to leverage the latest developments in AI. Demoing even the v1 of stable diffusion to the non-technical general users blows them away completely. Now that v2 is here, it’s clear we’re not able to keep pace in developing products to take advantage of it. The general public still is blown away by autosuggest in mobile OS keyboards. Very few really know how far…

While i agree it is exciting, the media industry will remain the same in size. Does this have applications outside media/entertainment ?

Re: Stable Diffusion 2.0

#390
Well darn. This is an awesome leap, but I've spent the last few months making a card game using Stable Diffusion art and I guess now I need to go back and go over everything again. Congratulations to the SD team on another wonderful step forward!
Post reply on HN