Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

181–190 of 661 posts

Re: Imagen, a text-to-image diffusion model

#181

Earlier quoted context omitted.

It depends on whether you'd like the model to learn casual or correlative relationships. If you want the model to understand what a "nurse" actually is, then it shouldn't be associated with female. If you want the model to understand how the word "nurse" is usually used, without regard for what a "nurse" actually is, then associating it with female is fine. The issue with a correlative model is that it can easily be…

At the end of a day, if you ask for a nurse, should the model output a male or female by default? If the input text lacks context/nuance, then the model must have some bias to infer the user's intent. This holds true for any image it generates; not just the politically sensitive ones. For example, if I ask for a picture of a person, and don't get one with pink hair, is that a shortcoming of the model? I'd say that bi…

> At the end of a day, if you ask for a nurse, should the model output a male or female by default?

Randomly pick one.

> Trying to generate a model that's "free of correlative relationships" is impossible because the model would never have the infinitely pedantic input text to describe the exact output image.

Sure, and you can never make a medical procedure 100% safe. Doesn't mean that you don't try to make them safer. You can trim the obvious low hanging fruit though.

Re: Imagen, a text-to-image diffusion model

#182
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

> a tendency for images portraying different professions to align with Western gender stereotypes

There are two possible ways of interpreting interpreting "gender stereotypes in professions".

biased or correct

https://www.abc.net.au/news/2018-05-21/the-most-gendered-top...

https://www.statista.com/statistics/1019841/female-physician...

Re: Imagen, a text-to-image diffusion model

#183

Earlier quoted context omitted.

Translation: we need to hand-tune this to not reflect reality but instead the world as we (Caucasian/Asian male American woke upper-middle class San Fransisco engineers) wish it to be. Maybe that's a nice thing, I wouldn't say their values are wrong but let's call a spade a spade.

"Reality" as defined by the available training set isn't necessarily reality. For example, Google's image search results pre-tweaking had some interesting thoughts on what constitutes a professional hairstyle, and that searches for "men" and "women" should only return light-skinned people: https://www.theguardian.com/technology/2016/apr/08/does-goog... Does that reflect reality? No. (I suspect there are also mostly u…

unstated but very real concerns

I say let people generate their own reality. The sooner the masses realise that ceci n'est pas une pipe , the less likely they are to be swayed by the growing un-reality created by companies like Google.

Re: Imagen, a text-to-image diffusion model

#185
post #42

Earlier quoted context omitted.

This raises some really interesting questions. We certainly don't want to perpetuate harmful stereotypes. But is it a flaw that the model encodes the world as it really is, statistically, rather than as we would like it to be? By this I mean that there are more light-skinned people in the west than dark, and there are more women nurses than men, which is reflected in the model's training data. If the model only gener…

> But is it a flaw that the model encodes the world as it really is Does a bias towards lighter skin represent reality? I was under the impression that Caucasians are a minority globally. I read the disclaimer as "the model does NOT represent reality".

Caucasians are overrepresented in internet pictures.

Re: Imagen, a text-to-image diffusion model

#186
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

Translation: we need to hand-tune this to not reflect reality but instead the world as we (Caucasian/Asian male American woke upper-middle class San Fransisco engineers) wish it to be. Maybe that's a nice thing, I wouldn't say their values are wrong but let's call a spade a spade.

"As we wish it to be" is not totally true, because there are some places where humanity's iconographic reality (which Imagen trains on) differs significantly from actual reality.

One example would be if Imagen draws a group of mostly white people when you say "draw a group of people". This doesn't reflect actual reality. Another would be if Imagen draws a group of men when you say "draw a group of doctors".

In these cases where iconographic reality differs from actual reality, hand-tuning could be used to bring it closer to the real world, not just the world as we might wish it to be!

I agree there's a problem here. But I'd state it more as "new technologies are being held to a vastly higher standard than existing ones." Imagine TV studios issuing a moratorium on any new shows that made being white (or rich) seem more normal than it was! The public might rightly expect studios to turn the dials away from the blatant biases of the past, but even if this would be beneficial the progressive and activist public is generations away from expecting a TV studio to not release shows until they're confirmed to be bias-free.

That said, Google's decision to not publish is probably less about the inequities in AI's representation of reality and more about the AI sometimes spitting out drawings that are offensive in the US, like racist caricatures.

Re: Imagen, a text-to-image diffusion model

#188
post #95

Earlier quoted context omitted.

It depends on whether you'd like the model to learn casual or correlative relationships. If you want the model to understand what a "nurse" actually is, then it shouldn't be associated with female. If you want the model to understand how the word "nurse" is usually used, without regard for what a "nurse" actually is, then associating it with female is fine. The issue with a correlative model is that it can easily be…

Additionally, if you optimize for most-likely-as-best, you will end up with the stereotypical result 100% of the time, instead of in proportional frequency to the statistics. Put another way, when we ask for an output optimized for "nursiness", is that not a request for some ur stereotypical nurse?

You could stipulate that it roll a die based on percentage results - if 70% of Americans are "white", then 70% of the time show a white person - 13% of the time the result should be black, etc.

That's excessively simplified but wouldn't this drop the stereotype and better reflect reality?

Re: Imagen, a text-to-image diffusion model

#189
post #143
post #43

Earlier quoted context omitted.

I know you're anon trolling, but the authors' names are: Chitwan Saharia, William Chan, Saurabh Saxena†, Lala Li†, Jay Whang†, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho†, David Fleet†, Mohammad Norouzi

Absolutely not related to the whole discussion, but what do "†" stands for?

https://en.wikipedia.org/wiki/Dagger_(mark)

Re: Imagen, a text-to-image diffusion model

#190
post #117
post #89

Earlier quoted context omitted.

You mean one of Google's domains? # whois appspot.com [Querying whois.verisign-grs.com] [Redirected to whois.markmonitor.com] [Querying whois.markmonitor.com] [whois.markmonitor.com] Domain Name: appspot.com Registry Domain ID: 145702338_DOMAIN_COM-VRSN Registrar WHOIS Server: whois.markmonitor.com Registrar URL: http://www.markmonitor.com Updated Date: 2022-02-06T09:29:56+0000 Creation Date: 2005-03-10T02:27:55+0000…

While appspot.com is a Google domain, anyone can register domains under it. It would be similarly surprising to see an official GitHub blog post under someproject.github.io

You mean like: https://say-can.github.io/

This is common in the research PA. People don't want to deal with broccoli man [1].

[1] https://www.youtube.com/watch?v=3t6L-FlfeaI

Post reply on HN