Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

221–230 of 661 posts

Re: Imagen, a text-to-image diffusion model

#221
post #159
post #74

Earlier quoted context omitted.

I don't see how this gets us (much) closer to general AI. Where is the reasoning?

Perhaps the confluence of NLP and something generative?

Yes metaculus mostly bet a magic number based on perhaps and tbh why not, the interaction of NLP and vision is mysterious and has potential. However those magic numbers should still be considered magic numbers. I agree that in 2040 the interactions will have extensively been studied though but the conclusion of wether we czn go much further on cross-models synergies is totally unknown or pessimist.

Re: Imagen, a text-to-image diffusion model

#222
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

it isn't woke enough. Lol.

In discussions like this, I always head for the gray-text comments to enjoy the last crumbs of the common sense in this world.

Re: Imagen, a text-to-image diffusion model

#223
post #74

Earlier quoted context omitted.

I don't see how this gets us (much) closer to general AI. Where is the reasoning?

Big pretrained models are good enough now that we can pipe them together in really cool ways and our representations of text and images seem to capture what we “mean.”

Yeah, it seems like it. But it's still just complicated statistical models. Again, where is the reasoning?

Re: Imagen, a text-to-image diffusion model

#224

Earlier quoted context omitted.

Yes, the idea is that just because it doesn't align to Western ideals of what seems unbiased doesn't mean that the same is necessarily true for other cultures, and by failing to release the model because it doesn't conform to Western, left wing cultural expectations, the authors are ignoring the diversity of cultures that exist globally.

No, it's coming from a perspective of moral realism. It's an objective moral truth that racial and ethnic biases are bad. Yet most cultures around the world are racist to at least some degree, and to they extent that the cultures do, they are bad. The argument you're making, paraphrased, is that the idea that biases are bad is itself situated in particular cultural norms. While that is true to some degree, from a mor…

You're confused by the double meaning of the word "bias".

Here we mean mathematical biases.

For example, a good mathematical model will correctly tell you that people in Japan (geographical term) are more likely to be Japanese (ethnic / racial bias). That's not "objectively morally bad", but instead, it's "correct".

Re: Imagen, a text-to-image diffusion model

#225
post #70

I give it a few years before Google makes stock images irrelevant.

The entire "content" industry could get eaten by a few hundred people curating + touching-up output from these models.

No, competitive advantage means that it’s impossible to run out of jobs just because someone/something is better at it than you.

(Consumer demand and boredom both being infinite is another thing working against it.)

Re: Imagen, a text-to-image diffusion model

#226
post #42
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

This raises some really interesting questions. We certainly don't want to perpetuate harmful stereotypes. But is it a flaw that the model encodes the world as it really is, statistically, rather than as we would like it to be? By this I mean that there are more light-skinned people in the west than dark, and there are more women nurses than men, which is reflected in the model's training data. If the model only gener…

Yes, there is a denominator problem. When selecting a sample "at random," what do you want the denominator to be? It could be "people in the US", "people in the West" (whatever countries you mean by that) or "people worldwide."

Also, getting a random sample of any demographic would be really hard, so no machine learning project is going to do that. Instead you've got a random sample of some arbitrary dataset that's not directly relevant to any particular purpose.

This is, in essence, a design or artistic problem: the Google researchers have some idea of what they want the statistical properties of their image generator to look like. What it does isn't it. So, artistically, the result doesn't meet their standards, and they're going to fix it.

There is no objective, universal, scientifically correct answer about which fictional images to generate. That doesn't all art is equally good, or that you should just ship anything without looking at quality along various axes.

Re: Imagen, a text-to-image diffusion model

#227

Earlier quoted context omitted.

At the end of a day, if you ask for a nurse, should the model output a male or female by default? If the input text lacks context/nuance, then the model must have some bias to infer the user's intent. This holds true for any image it generates; not just the politically sensitive ones. For example, if I ask for a picture of a person, and don't get one with pink hair, is that a shortcoming of the model? I'd say that bi…

> At the end of a day, if you ask for a nurse, should the model output a male or female by default? Randomly pick one. > Trying to generate a model that's "free of correlative relationships" is impossible because the model would never have the infinitely pedantic input text to describe the exact output image. Sure, and you can never make a medical procedure 100% safe. Doesn't mean that you don't try to make them safe…

> Randomly pick one.

How does the model back out the "certain people would like to pretend it's a fair coin toss that a randomly selected nurse is male or female" feature?

It won't be in any representative training set, so you're back to fishing for stock photos on getty rather than generating things.

Re: Imagen, a text-to-image diffusion model

#228

Earlier quoted context omitted.

Good lord. Withheld? They've published their research, they just aren't making the model available immediately, waiting until they can re-implement it so that you don't get racial slurs popping up when you ask for a cup of "black coffee." >While a subset of our training data was filtered to removed noise and undesirable content, such as pornographic imagery and toxic language, we also utilized LAION-400M dataset whic…

I wonder why they don't like the idea of autogenerated porn... They're already putting most artists out of a job, why not put porn stars out of a job too?

There's definitely a market for autogenerated porn. But automated porn in a Google branded model for general use around stuff that isn't necessarily intended to be pornographic, on the other hand...

Re: Imagen, a text-to-image diffusion model

#229

Earlier quoted context omitted.

Good lord. Withheld? They've published their research, they just aren't making the model available immediately, waiting until they can re-implement it so that you don't get racial slurs popping up when you ask for a cup of "black coffee." >While a subset of our training data was filtered to removed noise and undesirable content, such as pornographic imagery and toxic language, we also utilized LAION-400M dataset whic…

I wonder why they don't like the idea of autogenerated porn... They're already putting most artists out of a job, why not put porn stars out of a job too?

Copenhagen ethics (used by most people) require that all negative outcomes of a thing X become yours if you interact with X. It is not sensible to interact with high negativity things unless you are single-issue. It is logical for Google to not attempt to interact with porn where possible.

Re: Imagen, a text-to-image diffusion model

#230
post #42

Earlier quoted context omitted.

This raises some really interesting questions. We certainly don't want to perpetuate harmful stereotypes. But is it a flaw that the model encodes the world as it really is, statistically, rather than as we would like it to be? By this I mean that there are more light-skinned people in the west than dark, and there are more women nurses than men, which is reflected in the model's training data. If the model only gener…

It’s the same as with an artist: “hey artist, draw me a nurse.” “Hmm okay, do you want it a guy or girl?” “Don’t ask me, just draw what I’m saying.” The artist can then say: “Okay, but accept my biases.” or “I can’t since your input is ambiguous.” For a one-shot generative algorithm you must accept the artist’s biases.

Revert back to average representation of a nurse (give no weight to unspecified criterias, gender, age, skin-color, religion, country, hair-style, no style whether it's a drawing or a photography, no information about the year it was made, etc).

“hey artist, draw me a nurse.”

“Hmm okay, do you want it a guy or girl?”

“Don’t ask me, just draw what I’m saying.”

- Ok, I'll draw you what an average nurse looks like.

- Wait, it's a woman! She wears a nurse blouse and she has a nurse cap.

- Is it bad ?

- No.

- Ok then what's the problem, you asked for something that looked like a nurse but didn't specify anything else ?

Post reply on HN