Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

141–150 of 661 posts

Re: Imagen, a text-to-image diffusion model

#141
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

Indeed. If a project has shortcomings, why not just acknowledge the shortcomings and plan to improve on them in a future release? Is it anticipated that "engineer" being rendered as a man by the model is going to be an actively dangerous thing to have out in the world?

"what could go wrong anyway?"

Re: Imagen, a text-to-image diffusion model

#143
post #43

Earlier quoted context omitted.

Translation: we need to hand-tune this to not reflect reality but instead the world as we (Caucasian/Asian male American woke upper-middle class San Fransisco engineers) wish it to be. Maybe that's a nice thing, I wouldn't say their values are wrong but let's call a spade a spade.

I know you're anon trolling, but the authors' names are: Chitwan Saharia, William Chan, Saurabh Saxena†, Lala Li†, Jay Whang†, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho†, David Fleet†, Mohammad Norouzi

Absolutely not related to the whole discussion, but what do "†" stands for?

Re: Imagen, a text-to-image diffusion model

#144
post #80
post #42

Earlier quoted context omitted.

This raises some really interesting questions. We certainly don't want to perpetuate harmful stereotypes. But is it a flaw that the model encodes the world as it really is, statistically, rather than as we would like it to be? By this I mean that there are more light-skinned people in the west than dark, and there are more women nurses than men, which is reflected in the model's training data. If the model only gener…

I think the statistics/representation problem is a big problem on its own, but IMO the bigger problem here is democratizing access to human-like creativity. Currently, the ability to create compelling art is only held by those with some artistic talent. With a tool like this, that restriction is gone. Everyone, no matter how uncreative, untalented, or uncommitted, can create compelling visuals, provided they can use…

I can't quite tell if you're being sarcastic about people being able to make things other people would find offensive being a problem. Are you missing an /s?

Re: Imagen, a text-to-image diffusion model

#145
post #62

Earlier quoted context omitted.

If your query was about hairstyle, why do you even look or care about the skin color ? Nowhere there is any precision for a preferred skin color in the query of th user. So it sorts and gives the most average examples based on the examples that were found on the internet. Essentially answering the query "SELECT * FROM `non-professional hairstyles` ORDER BY score DESC LIMIT 10". It's like if you search on Google "best…

> If your query was about hairstyle, why do you even look at the skin color ? You know that race has a large effect on hair right?

I'd be careful where you're going with that. You might make a point that is the opposite of what you intended.

Re: Imagen, a text-to-image diffusion model

#146
post #41

Really impressive. If we are able to generate such detailed images, is there anything similar for text to music? I would I though that it would be simpler to achieve than text to image.

why stop at audio? the pinnacle of this would be text-to-videos, equally indistinguishable from real thing.

The way things look when still is much easier to fake than the way things move.

I would expect AI development to follow a similar path to digital media generally, as its following the increasing difficulty and space requirements of digitally representing said media: text What’s more impressive to me is how far ahead text-to-speech is, but I think the explanation is straightforward (the accessibility value has motivated us to work on that for a lot longer).

Re: Imagen, a text-to-image diffusion model

#147

Why is this seemingly official Google blog post on this random non-Google domain?

This is quite suspicious considering that google AI research has an official blog[1], and this is not mentioned at all there. It seems quite possible that this is an elaborate prank.

1: https://ai.googleblog.com/

Re: Imagen, a text-to-image diffusion model

#149

Earlier quoted context omitted.

> it will be used in ways we haven’t anticipated Oh yeah, as a woman who grew up in a Third World country, how an AI model generates images would have deeply affected my daily struggles! /s It's kinda insulting that they think that this would be insulting. Like "Oh no I asked the model to draw a doctor and it drew a male doctor, I guess there's no point in me pursuing medical studies" ...

Yes actually, subconscious bias due to historical prejudice does have a large effect on society. Obviously there are things with much larger effects, that doesn't mean that this doesn't exist. > Oh no I asked the model to draw a doctor and it drew a male doctor, I guess there's no point in me pursuing medical studies If you don't think this is a real thing that happens to children you're not thinking especially hard.…

> If you don't think this is a real thing that happens to children you're not thinking especially hard

I believe that's where parenting comes in. Maybe I'm too cynical but I think that the parents' job is to undo all of the harm done by society and instill in their children the "correct" values.

Re: Imagen, a text-to-image diffusion model

#150

Earlier quoted context omitted.

It depends on whether you'd like the model to learn casual or correlative relationships. If you want the model to understand what a "nurse" actually is, then it shouldn't be associated with female. If you want the model to understand how the word "nurse" is usually used, without regard for what a "nurse" actually is, then associating it with female is fine. The issue with a correlative model is that it can easily be…

> If you want the model to understand how the word "nurse" is usually used, without regard for what a "nurse" actually is, then associating it with female is fine. That’s a distinction without a difference. Meaning is use.

Not really; the gender of a nurse is accidental, other properties are essential.
Post reply on HN