Live data from Hacker News

Direct pixel-space megapixel image generation with diffusion models

crowsonkb.github.io

41–50 of 50 posts

Re: Direct pixel-space megapixel image generation with diffusion models

#41
post #7

Earlier quoted context omitted.

Which discord if its open to the public? I was on one woth kath in 2021 and loved her insights, would love to again

Same; a good ML focused discord would be great. Training ViTs all day is lonely work. I'm mostly locked into skimming the "Research" channels of image generation discords. LAION used to be decent with a good amount of interesting discussion, but it seems to have devolved into toxicity in the last year.

See my other comment replying to that.

Re: Direct pixel-space megapixel image generation with diffusion models

#42

I'm one of the authors; happy to answer questions. this arch is of course nice for high-resolution synthesis, but there's some other cool stuff worth mentioning.. activations are small! so you can enjoy bigger batch sizes. this is due to the 4x patching we do on the ingress to the model, and the effectiveness of neighbourhood attention in joining patches at the seams. the model's inductive biases are pretty different…

[flagged]

Re: Direct pixel-space megapixel image generation with diffusion models

#43

I'm one of the authors; happy to answer questions. this arch is of course nice for high-resolution synthesis, but there's some other cool stuff worth mentioning.. activations are small! so you can enjoy bigger batch sizes. this is due to the 4x patching we do on the ingress to the model, and the effectiveness of neighbourhood attention in joining patches at the seams. the model's inductive biases are pretty different…

Did you do any inpainting experiments? I can imagine a pixel-space diffusion model to be better at it than one with a latent auto-encoder.

Re: Direct pixel-space megapixel image generation with diffusion models

#44
post #43

I'm one of the authors; happy to answer questions. this arch is of course nice for high-resolution synthesis, but there's some other cool stuff worth mentioning.. activations are small! so you can enjoy bigger batch sizes. this is due to the 4x patching we do on the ingress to the model, and the effectiveness of neighbourhood attention in joining patches at the seams. the model's inductive biases are pretty different…

Did you do any inpainting experiments? I can imagine a pixel-space diffusion model to be better at it than one with a latent auto-encoder.

Not yet, we focused on the architecture for this paper. I totally agree with you though - pixel space is generally less limiting than a latent space for diffusion, so we would expect good performance inpainting behavior and other editing tasks.

Re: Direct pixel-space megapixel image generation with diffusion models

#45

I'm one of the authors; happy to answer questions. this arch is of course nice for high-resolution synthesis, but there's some other cool stuff worth mentioning.. activations are small! so you can enjoy bigger batch sizes. this is due to the 4x patching we do on the ingress to the model, and the effectiveness of neighbourhood attention in joining patches at the seams. the model's inductive biases are pretty different…

I appreciate the restraint of showing the speedup on a log-scale chart rather than trying to show a 99% speed up any other way.

I see your headline speed comparison is to "Pixel-space DiT-B/4" - but how does your model compare to the likes of SDXL? I gather they spent $$$$$$ on training etc, so I'd understand if direct comparisons don't make sense.

And do you have any results on things that are traditionally challenging for generative AI, like clocks and mirrors?

Re: Direct pixel-space megapixel image generation with diffusion models

#48

Can I ask one basic thing --> From what are the images generated?

Not text but an id representing a class/category of images from the dataset. Or they are “unconditional” and the model tries to output something similar to a random image from the dataset each time.

Re: Direct pixel-space megapixel image generation with diffusion models

#49

Can I ask one basic thing --> From what are the images generated?

The models presented in the paper are trained on class-conditional ImageNet (where the input is Gaussian noise and one of 1000 classes, e.g., "car") and unconditional FFHQ (where the input is only Gaussian noise).

Re: Direct pixel-space megapixel image generation with diffusion models

#50
post #5

Earlier quoted context omitted.

Making a bank account work for you is a hard discipline and requires budgeting and the like. True, not all of us can "hack" it, but that doesn't mean that with some community classes and help you'll be able to use your bank account well! Why do I feel like a chatbot with this message.

You're actually talking to a bot, in this particular case. 12 minutes old with -2 karma. :berk:

Accusations of being a bot are explicitly against the rules.
Post reply on HN