Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

251–260 of 661 posts

Re: Imagen, a text-to-image diffusion model

#251

As someone who has a layman's understanding of neural networks, and who did some neural network programming ~20 years ago before the real explosion of the field, can someone point to some resources where I can get a better understanding about how this magic works? I mean, from my perspective, the skill in these (and DALL-E's) image reproductions is truly astonishing. Just looking for more information about how the so…

Figure A.4 in the linked paper is a good high level overview of this model. Shame it was hidden away on page 19 in the appendix!

Each box you see there has a section in the paper explaining it in more detail.

Re: Imagen, a text-to-image diffusion model

#253

As someone who has a layman's understanding of neural networks, and who did some neural network programming ~20 years ago before the real explosion of the field, can someone point to some resources where I can get a better understanding about how this magic works? I mean, from my perspective, the skill in these (and DALL-E's) image reproductions is truly astonishing. Just looking for more information about how the so…

Check https://github.com/multimodalart/majesty-diffusion or https://github.com/lucidrains/DALLE2-pytorch

There is a Google Colab workbook that you can try and run for free :)

This is the image-text pairs behind: https://laion.ai/laion-400-open-dataset/

Re: Imagen, a text-to-image diffusion model

#254

Earlier quoted context omitted.

> But is it a flaw that the model encodes the world as it really is Does a bias towards lighter skin represent reality? I was under the impression that Caucasians are a minority globally. I read the disclaimer as "the model does NOT represent reality".

Caucasians are overrepresented in internet pictures.

This, I would imagine this heavily correlates to things like income and gdp per capita.

Re: Imagen, a text-to-image diffusion model

#255
post #136

Earlier quoted context omitted.

> Every dataset is biased. Sure, I wasn't questioning the bias of the data, I was talking about the bias of the real world and whether we want the model to be "unbiased about bias" i.e. metabiased or not. Showing nurses equally as men and women is not biased, but it's metabiased, because the real world is biased. Whether metabias is right or not is more interesting than the question of whether bias is wrong because i…

Please be kinder to yourself. You need to be your own strongest advocate, and that's not incompatible with being humble. You have plenty to contribute to this world, and the vast majority of us appreciate what you have to offer.

Agreed. They are valid points clearly stated and a valuable contribution to the discussion.

Re: Imagen, a text-to-image diffusion model

#256
post #223

Earlier quoted context omitted.

Big pretrained models are good enough now that we can pipe them together in really cool ways and our representations of text and images seem to capture what we “mean.”

Yeah, it seems like it. But it's still just complicated statistical models. Again, where is the reasoning?

I don’t care whether it reasons its way from “3 teddy bears below 7 flamingos” to a picture of that or if it gets there some other way.

But also, some of the magic in having good enough pretrained representations is that you don’t need to train them further for downstream tasks, which means non-differentiable tasks like logic could soon become more tenable.

Re: Imagen, a text-to-image diffusion model

#257
post #42
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

This raises some really interesting questions. We certainly don't want to perpetuate harmful stereotypes. But is it a flaw that the model encodes the world as it really is, statistically, rather than as we would like it to be? By this I mean that there are more light-skinned people in the west than dark, and there are more women nurses than men, which is reflected in the model's training data. If the model only gener…

This sounds like descriptivism vs prescriptivism. In English (native language) I’m a descriptivist, in all other languages I have to tell myself to be a prescriptivist while I’m actively learning and then switch back to descriptivism to notice when the lessons were wrong or misleading.

Re: Imagen, a text-to-image diffusion model

#258
post #150

Earlier quoted context omitted.

Not really; the gender of a nurse is accidental, other properties are essential.

While not essential, I wouldn't exactly call the gender "accidental": > We investigated sex differences in 473,260 adolescents’ aspirations to work in things-oriented (e.g., mechanic), people-oriented (e.g., nurse), and STEM (e.g., mathematician) careers across 80 countries and economic regions using the 2018 Programme for International Student Assessment (PISA). We analyzed student career aspirations in combination…

The "Gender Equality Paradox"... there's a fascinating episode[0] about it. It's incredible how unscientific and ideologically-motivated one side comes off in it.

0. https://www.youtube.com/watch?v=_XsEsTvfT-M

Re: Imagen, a text-to-image diffusion model

#259
post #150

Earlier quoted context omitted.

> If you want the model to understand how the word "nurse" is usually used, without regard for what a "nurse" actually is, then associating it with female is fine. That’s a distinction without a difference. Meaning is use.

Not really; the gender of a nurse is accidental, other properties are essential.

How do you know this? Because you can, in your mind, divide the function of a nurse from the statistical reality of nursing?

Are the logical divisions you make in your mind really indicative of anything other than your arbitrary personal preferences?

Re: Imagen, a text-to-image diffusion model

#260

Earlier quoted context omitted.

Copenhagen ethics (used by most people) require that all negative outcomes of a thing X become yours if you interact with X. It is not sensible to interact with high negativity things unless you are single-issue. It is logical for Google to not attempt to interact with porn where possible.

> Copenhagen ethics (used by most people) The idea that most people use any coherent ethical framework (even something as high level and nearly content-free as Copenhagen) much less a particular coherent ethical framework is, well, not well supported by the evidence. > require that all negative outcomes of a thing X become yours if you interact with X. It is not sensible to interact with high negativity things unless…

I'm sure you are capable of steelmanning the argument.
Post reply on HN