Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

191–200 of 661 posts

Re: Imagen, a text-to-image diffusion model

#191

Earlier quoted context omitted.

Except "reality" in this case is just their biased training set. E.g. There's more non-white doctors and nurses in the world than white ones, yet their model would likely show an image of white person when you type in "doctor".

Alternately, there are more females nurses in the world than male nurses, and their model probably shows an image of a woman when you type in "nurse" but they consider that a problem.

Google Image Search doesn’t reflect harsh reality when you search for things; it shows you what’s on Pinterest. The same is more likely to apply here than the idea they’re trying to hide something.

There’s no reason to believe their model training learns the same statistics as their input dataset even. If that’s not an explicit training goal then whatever happens happens. AI isn’t magic or more correct than people.

Re: Imagen, a text-to-image diffusion model

#192
post #136

Earlier quoted context omitted.

> But is it a flaw that the model encodes the world as it really is I want to be clear here, bias can be introduced at many different points. There's dataset bias, model bias, and training bias. Every model is biased. Every dataset is biased. Yes, the real world is also biased. But I want to make sure that there are ways to resolve this issue. It is terribly difficult, especially in a DL framework (even more so in a…

> Every dataset is biased. Sure, I wasn't questioning the bias of the data, I was talking about the bias of the real world and whether we want the model to be "unbiased about bias" i.e. metabiased or not. Showing nurses equally as men and women is not biased, but it's metabiased, because the real world is biased. Whether metabias is right or not is more interesting than the question of whether bias is wrong because i…

Please be kinder to yourself. You need to be your own strongest advocate, and that's not incompatible with being humble. You have plenty to contribute to this world, and the vast majority of us appreciate what you have to offer.

Re: Imagen, a text-to-image diffusion model

#193

Earlier quoted context omitted.

It depends on whether you'd like the model to learn casual or correlative relationships. If you want the model to understand what a "nurse" actually is, then it shouldn't be associated with female. If you want the model to understand how the word "nurse" is usually used, without regard for what a "nurse" actually is, then associating it with female is fine. The issue with a correlative model is that it can easily be…

At the end of a day, if you ask for a nurse, should the model output a male or female by default? If the input text lacks context/nuance, then the model must have some bias to infer the user's intent. This holds true for any image it generates; not just the politically sensitive ones. For example, if I ask for a picture of a person, and don't get one with pink hair, is that a shortcoming of the model? I'd say that bi…

This type of bias sounds a lot easier to explain away as a non-issue when we are using "nurse" as the hypothetical prompt. What if the prompt is "criminal", "rapist", or some other negative? Would that change your thought process or would you be okay with the system always returning a person of the same race and gender that statistics indicate is the most likely? Do you see how that could be a problem?

Re: Imagen, a text-to-image diffusion model

#194
post #85

Earlier quoted context omitted.

If you type as a prompt "most beautiful woman in the world", you get a brown-skinned brown-haired woman with hazel eyes. What should be the right answer then ? You put a blonde, you offend the brown haired. You put blue eyes, you offend the brown eyes. etc.

That's an unanswerable question. Perhaps the answer is "don't". Siri takes this approach for a wide range of queries.

I think the key is to take the information in this world with a little bit pinch of salt.

When you do a search on a search engine, the results are biased too, but still, they shouldn't be artificially censored to fit some political views.

I asked one algorithm few minutes ago (it's called t0pp and it's free to try online, and it's quite fascinating because it's uncensored):

"What is the name of the most beautiful man on Earth ?

- He is called Brad Pitt."

==

Is it true in an objective way ? Probably not.

Is there an actual answer ? Probably yes, there is somewhere a man who scores better than the others.

Is it socially acceptable ? Probably not.

The question is:

If you interviewed 100 persons in the street, and asked the question "What is the name of the most beautiful man on Earth ?".

I'm pretty sure you'd get Brad Pitt often coming in.

Now, what about China ?

We don't have many examples there, they have no clue who is Brad Pitt probably, and there is probably someone else that is considered more beautiful by over 1B people

(t0pp tells me it's someone called "Zhu Zhu" :D )

==

Two solutions:

1) Censorship

-> Sorry there is too much bias in Western and we don't want to offend anyone, no answer, or a generic overriding human answer that is safe for advertisers, but totally useless ("the most beautiful human is you")

2) Adding more examples

-> Work on adding more examples from abroad trying to get the "average human answer".

==

I really prefer solution (2) in the core algorithms and dataset development, rather than going through (1).

(1) is more a choice to make at the stage when you are developing a virtual psychologist or a chat assistant, not when creating AI building blocks.

Re: Imagen, a text-to-image diffusion model

#195
This competitor might be better for respecting spatial prepositions and photorealism but on a quick look i find the images more uncanny. DALL-E has IMHO better camera POV/distance and is able to make artistic/dreamy/beautiful images. I haven't yet seen this Google model be competitive for art and uncaniness. However progress is great and I might be wrong.

Re: Imagen, a text-to-image diffusion model

#196
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

Yup this is what happens when people who want headlines nitpick for bullshit in a state-of-the-art model which simply reflects the state of the society. Better not to release the model itself than keep explaining over and over how a model is never perfect.

Re: Imagen, a text-to-image diffusion model

#197
post #150

Earlier quoted context omitted.

> If you want the model to understand how the word "nurse" is usually used, without regard for what a "nurse" actually is, then associating it with female is fine. That’s a distinction without a difference. Meaning is use.

Not really; the gender of a nurse is accidental, other properties are essential.

While not essential, I wouldn't exactly call the gender "accidental":

> We investigated sex differences in 473,260 adolescents’ aspirations to work in things-oriented (e.g., mechanic), people-oriented (e.g., nurse), and STEM (e.g., mathematician) careers across 80 countries and economic regions using the 2018 Programme for International Student Assessment (PISA). We analyzed student career aspirations in combination with student achievement in mathematics, reading, and science, as well as parental occupations and family wealth. In each country and region, more boys than girls aspired to a things-oriented or STEM occupation and more girls than boys to a people-oriented occupation. These sex differences were larger in countries with a higher level of women's empowerment. We explain this counter-intuitive finding through the indirect effect of wealth. Women's empowerment is associated with relatively high levels of national wealth and this wealth allows more students to aspire to occupations they are intrinsically interested in.

Source: https://psyarxiv.com/zhvre/ (HN discussion: https://news.ycombinator.com/item?id=29040132)

Re: Imagen, a text-to-image diffusion model

#198
post #59
post #51

Earlier quoted context omitted.

Looks like no, "The potential risks of misuse raise concerns regarding responsible open-sourcing of code and demos. At this time we have decided not to release code or a public demo. In future work we will explore a framework for responsible externalization that balances the value of external auditing with the risks of unrestricted open-access."

> the risks of unrestricted open-access What exactly is the risk?

If the model is used to generate offensive imagery, it may result in a negative press response directed at the company.

Re: Imagen, a text-to-image diffusion model

#199
post #143
post #43

Earlier quoted context omitted.

I know you're anon trolling, but the authors' names are: Chitwan Saharia, William Chan, Saurabh Saxena†, Lala Li†, Jay Whang†, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho†, David Fleet†, Mohammad Norouzi

Absolutely not related to the whole discussion, but what do "†" stands for?

It's just a different asterisk to distinguish, in this case, in the paper, they are "core contributors."

Re: Imagen, a text-to-image diffusion model

#200
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

It’s wild to me that the HN consensus is so often that 1) discourse around the internet is terrible, it’s full of spam and crap, and the internet is an awful unrepresentative snapshot of human existence, and 2) the biases of general-internet-training-data are fine in ML models because it just reflects real life.

The bias on HN is that people who prioritize being nice, or may possibly have humanities degrees or be ultra-libs from SF, are wrong because the correct answer would be cynical and cold-heartedly mechanical.

Other STEM adjacent communities feel similarly but I don’t get it from actual in person engineers much.

Post reply on HN