Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

261–270 of 661 posts

Re: Imagen, a text-to-image diffusion model

#261

Earlier quoted context omitted.

Yes actually, subconscious bias due to historical prejudice does have a large effect on society. Obviously there are things with much larger effects, that doesn't mean that this doesn't exist. > Oh no I asked the model to draw a doctor and it drew a male doctor, I guess there's no point in me pursuing medical studies If you don't think this is a real thing that happens to children you're not thinking especially hard.…

> If you don't think this is a real thing that happens to children you're not thinking especially hard I believe that's where parenting comes in. Maybe I'm too cynical but I think that the parents' job is to undo all of the harm done by society and instill in their children the "correct" values.

> I think that the parents' job is to undo all of the harm done by society and instill in their children the "correct" values.

Far from being too cynical, this is too optimistic.

The vast majority of parents try to instill the value "do not use heroin." And yet society manages to do that harm on a large scale. There are other examples.

Re: Imagen, a text-to-image diffusion model

#262
post #43

Earlier quoted context omitted.

Translation: we need to hand-tune this to not reflect reality but instead the world as we (Caucasian/Asian male American woke upper-middle class San Fransisco engineers) wish it to be. Maybe that's a nice thing, I wouldn't say their values are wrong but let's call a spade a spade.

I know you're anon trolling, but the authors' names are: Chitwan Saharia, William Chan, Saurabh Saxena†, Lala Li†, Jay Whang†, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho†, David Fleet†, Mohammad Norouzi

[deleted]

Re: Imagen, a text-to-image diffusion model

#263
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

They're withholding the API, code, and trained data because they don't want it to affect their corporate image. The good thing is they released their paper which will allow easy reproduction. T5-XXL looks on par with CLIP so we may not see an open source version of T5 for a bit (LAION is working on reproducing CLIP), but this is all progress.

T5 was open-sourced on release (up to 11B params): https://github.com/google-research/text-to-text-transfer-tra...

It is also available via Hugging Face transformers.

However, the paper mentions T5-XXL is 4.6B, which doesn't fit any of the checkpoints above, so I'm confused.

Re: Imagen, a text-to-image diffusion model

#264
post #223

Earlier quoted context omitted.

Big pretrained models are good enough now that we can pipe them together in really cool ways and our representations of text and images seem to capture what we “mean.”

Yeah, it seems like it. But it's still just complicated statistical models. Again, where is the reasoning?

A belief oft shared is that sufficiently complicated statistical models are indistinguishable from reasoning.

Re: Imagen, a text-to-image diffusion model

#265

Earlier quoted context omitted.

> But is it a flaw that the model encodes the world as it really is Does a bias towards lighter skin represent reality? I was under the impression that Caucasians are a minority globally. I read the disclaimer as "the model does NOT represent reality".

Caucasians are overrepresented in internet pictures.

Right, that's the likely cause of the bias.

Re: Imagen, a text-to-image diffusion model

#266
post #179

Earlier quoted context omitted.

See section 6 titled “Conclusions, Limitations and Societal Impact” in the research paper: https://gweb-research-imagen.appspot.com/paper.pdf One quote: > “On the other hand, generative methods can be leveraged for malicious purposes, including harassment and misinformation spread [20], and raise many concerns regarding social and cultural exclusion and bias [67, 62, 68]”

But do we trust that those who do have access won't be using it for "malicious purposes" (which they might not think is malicious, but perhaps it is to those who don't have access)?

It's not up to you. It's up to them, and they trust themselves/don't care about your definition of malicious.

Re: Imagen, a text-to-image diffusion model

#267
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

From the HN rules:

>Eschew flamebait. Avoid unrelated controversies and generic tangents.

They provided a pretty thorough overview (nearly 500 words) of the multiple reasons why they are showing caution. You picked out the one that happened to bother you the most and have posted a misleading claim that the tech is being withheld entirely because of it.

Re: Imagen, a text-to-image diffusion model

#268
post #193

Earlier quoted context omitted.

At the end of a day, if you ask for a nurse, should the model output a male or female by default? If the input text lacks context/nuance, then the model must have some bias to infer the user's intent. This holds true for any image it generates; not just the politically sensitive ones. For example, if I ask for a picture of a person, and don't get one with pink hair, is that a shortcoming of the model? I'd say that bi…

This type of bias sounds a lot easier to explain away as a non-issue when we are using "nurse" as the hypothetical prompt. What if the prompt is "criminal", "rapist", or some other negative? Would that change your thought process or would you be okay with the system always returning a person of the same race and gender that statistics indicate is the most likely? Do you see how that could be a problem?

Cultural biases aren’t uniform across nations. If a prompt returns caucasians for nurses, and other races for criminals then most people in my country would not note that as racism simply because there are not, and there have never in history, been enough caucasians resident for anyone to create significant race theories about them.

This is a far cry from say the USA where that would instantly trigger a response since until the 1960s there was a widespread race based segregation.

Re: Imagen, a text-to-image diffusion model

#269

Earlier quoted context omitted.

> At the end of a day, if you ask for a nurse, should the model output a male or female by default? Randomly pick one. > Trying to generate a model that's "free of correlative relationships" is impossible because the model would never have the infinitely pedantic input text to describe the exact output image. Sure, and you can never make a medical procedure 100% safe. Doesn't mean that you don't try to make them safe…

what if I asked the model to show me a sunday school photograph of baptists in the National Baptist Convention?

The pictures I got from a similar model when asking for a "sunday school photograph of baptists in the National Baptist Convention": https://ibb.co/sHGZwh7

Re: Imagen, a text-to-image diffusion model

#270
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

It’s wild to me that the HN consensus is so often that 1) discourse around the internet is terrible, it’s full of spam and crap, and the internet is an awful unrepresentative snapshot of human existence, and 2) the biases of general-internet-training-data are fine in ML models because it just reflects real life.

Why is it wild? How is it contradictory?
Post reply on HN