Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

121–130 of 661 posts

Re: Imagen, a text-to-image diffusion model

#121

Earlier quoted context omitted.

Translation: we need to hand-tune this to not reflect reality but instead the world as we (Caucasian/Asian male American woke upper-middle class San Fransisco engineers) wish it to be. Maybe that's a nice thing, I wouldn't say their values are wrong but let's call a spade a spade.

Except "reality" in this case is just their biased training set. E.g. There's more non-white doctors and nurses in the world than white ones, yet their model would likely show an image of white person when you type in "doctor".

Alternately, there are more females nurses in the world than male nurses, and their model probably shows an image of a woman when you type in "nurse" but they consider that a problem.

Re: Imagen, a text-to-image diffusion model

#122
post #95

Earlier quoted context omitted.

It depends on whether you'd like the model to learn casual or correlative relationships. If you want the model to understand what a "nurse" actually is, then it shouldn't be associated with female. If you want the model to understand how the word "nurse" is usually used, without regard for what a "nurse" actually is, then associating it with female is fine. The issue with a correlative model is that it can easily be…

Additionally, if you optimize for most-likely-as-best, you will end up with the stereotypical result 100% of the time, instead of in proportional frequency to the statistics. Put another way, when we ask for an output optimized for "nursiness", is that not a request for some ur stereotypical nurse?

You could simply encode a score for how well the output matches the input. If 25% of trees in summer are brown, perhaps the output should also have 25% brown. The model scores itself on frequencies as well as correctness.

Re: Imagen, a text-to-image diffusion model

#123
post #62

Earlier quoted context omitted.

"Reality" as defined by the available training set isn't necessarily reality. For example, Google's image search results pre-tweaking had some interesting thoughts on what constitutes a professional hairstyle, and that searches for "men" and "women" should only return light-skinned people: https://www.theguardian.com/technology/2016/apr/08/does-goog... Does that reflect reality? No. (I suspect there are also mostly u…

If your query was about hairstyle, why do you even look or care about the skin color ? Nowhere there is any precision for a preferred skin color in the query of th user. So it sorts and gives the most average examples based on the examples that were found on the internet. Essentially answering the query "SELECT * FROM `non-professional hairstyles` ORDER BY score DESC LIMIT 10". It's like if you search on Google "best…

> If your query was about hairstyle, why do you even look at the skin color ?

You know that race has a large effect on hair right?

Re: Imagen, a text-to-image diffusion model

#124
post #29

Earlier quoted context omitted.

The ironic part is that these "social and cultural biases" are purely from a Western, American lens. The people writing that paragraph are completely oblivious to the idea that there could be other cultures other than the Western American one. In attempting to prevent "encoding of social and cultural biases" they have encoded such biases themselves into their own research.

It seems you've got it backwards: "tendency for images portraying different professions to align with Western gender stereotypes" means that they are calling out their own work precisely because it is skewed in the direction of Western American biases.

The very act of mentioning "western gender stereotypes" starts from a biased position.

Why couldn't they be "northern gender stereotypes"? Is the world best explained as a division of west/east instead of north/south? The northern hemisphere has much more population than the south, and almost all rich countries are in the northern hemisphere. And precisely it's these rich countries pushing the concept of gender stereotypes. In poor countries, nobody cares about these "gender stereotypes".

Actually, the lines dividing the earth into north and south, east and west hemispheres are arbitrary, so maybe they shouldn't mention the word "western" to avoid the propagation of stereotypes about earth regions.

Or why couldn't they be western age stereotypes? Why are there no kids or very old people depicted as nurses?

Why couldn't they be western body shape stereotypes? Why are there so few obese people in the images? Why are there no obese people depicted as athletes?

Are all of these really stereotypes or just natural consequences of natural differences?

Re: Imagen, a text-to-image diffusion model

#125
post #42

Earlier quoted context omitted.

This raises some really interesting questions. We certainly don't want to perpetuate harmful stereotypes. But is it a flaw that the model encodes the world as it really is, statistically, rather than as we would like it to be? By this I mean that there are more light-skinned people in the west than dark, and there are more women nurses than men, which is reflected in the model's training data. If the model only gener…

> But is it a flaw that the model encodes the world as it really is Does a bias towards lighter skin represent reality? I was under the impression that Caucasians are a minority globally. I read the disclaimer as "the model does NOT represent reality".

Well first, I didn't say caucasian; light-skinned includes Spanish people and many others that caucasian excludes, and that's why I said the former. Also, they are a minority globally, but the GP mentioned "Western stereotypes", and they're a majority in the West, so that's why I said "in the west" when I said that there are more light-skinned people.

Re: Imagen, a text-to-image diffusion model

#126
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

Good lord. Withheld? They've published their research, they just aren't making the model available immediately, waiting until they can re-implement it so that you don't get racial slurs popping up when you ask for a cup of "black coffee."

>While a subset of our training data was filtered to removed noise and undesirable content, such as pornographic imagery and toxic language, we also utilized LAION-400M dataset which is known to contain a wide range of inappropriate content including pornographic imagery, racist slurs, and harmful social stereotypes

Tossing that stuff when it comes up in a research environment is one thing, but Google clearly wants to implement this as a product, used all over the world by a huge range of people. If the dataset has problems, and why wouldn't it, it is perfectly rational to want to wait and re-implement it with a better one. DALL-E 2 was trained on a curated dataset so it couldn't generate sex or gore. Others are sanitizing their inputs too and have done for a long time. It is the only thing that makes sense for a company looking to commercialize a research project.

This has nothing to do with "inability to cope" and the implied woke mob yelling about some minor flaw. It's about building a tool that doesn't bake in serious and avoidable problems.

Re: Imagen, a text-to-image diffusion model

#127

Earlier quoted context omitted.

"Reality" as defined by the available training set isn't necessarily reality. For example, Google's image search results pre-tweaking had some interesting thoughts on what constitutes a professional hairstyle, and that searches for "men" and "women" should only return light-skinned people: https://www.theguardian.com/technology/2016/apr/08/does-goog... Does that reflect reality? No. (I suspect there are also mostly u…

The reality is that hair styles on the left side of the image in the article are widely considered unprofessional in today's workplaces. That may seem egregiously wrong to you, but it is a truth of American and European society today. Should it be Google's job to rewrite reality?

Only black people have unprofessional hair and only white people have professional hair is not reality.

Re: Imagen, a text-to-image diffusion model

#128
post #120

Earlier quoted context omitted.

appspot.com is the domain that hosts all App Engine apps (at least those that don't use a custom domain). It's kind of like Heroku and has been around for at least a decade. https://cloud.google.com/appengine

Spring 2008: 14 years!

Whoa, I feel super old, I first used it in 2011 when I thought it was new.

Re: Imagen, a text-to-image diffusion model

#129

All of these AI findings are cool in theory. But until its accessible to some decent amount of people/customers - its basically useless fluff. You can tell me those pictures are generated by an AI and I might believe it, but until real people can actually test it... it's easy enough to fake. This page isn't even the remotest bit legit by the URL, It looks nicely put together and that's about it. Could have easily put…

Inference times are key. If it can't be produced within reasonable latency, then there will be no real world use case for it because it's simply too expensive to run inference at scale.

Re: Imagen, a text-to-image diffusion model

#130

Earlier quoted context omitted.

Translation: we need to hand-tune this to not reflect reality but instead the world as we (Caucasian/Asian male American woke upper-middle class San Fransisco engineers) wish it to be. Maybe that's a nice thing, I wouldn't say their values are wrong but let's call a spade a spade.

Translation: we need to hand-tune this to not reflect reality Is it reflecting reality, though? Seems to me that (as with any ML stuff, right?) it's reflecting the training corpus. Futhermore, is it this thing's job to reflect reality? the world as we (Caucasian/Asian male American woke upper-middle class San Fransisco engineers) wish it to be Snarky answer: Ah, yes, let's make sure that things like "A giant cobra sn…

> Snarky answer: Ah, yes, let's make sure that things like "A giant cobra snake on a farm. The snake is made out of corn" reflect reality.

If it didn't reflect reality, you wouldn't be impressed by the image of the snake made of corn.

Post reply on HN