Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

231–240 of 661 posts

Re: Imagen, a text-to-image diffusion model

#231
post #193

Earlier quoted context omitted.

At the end of a day, if you ask for a nurse, should the model output a male or female by default? If the input text lacks context/nuance, then the model must have some bias to infer the user's intent. This holds true for any image it generates; not just the politically sensitive ones. For example, if I ask for a picture of a person, and don't get one with pink hair, is that a shortcoming of the model? I'd say that bi…

This type of bias sounds a lot easier to explain away as a non-issue when we are using "nurse" as the hypothetical prompt. What if the prompt is "criminal", "rapist", or some other negative? Would that change your thought process or would you be okay with the system always returning a person of the same race and gender that statistics indicate is the most likely? Do you see how that could be a problem?

Not the person you responded to, but I do see how someone could be hurt by that, and I want to avoid hurting people. But is this the level at which we should do it? Could skewing search results, i.e. hiding the bias of the real world, give us the impression that everything is fine and we don't need to do anything to actually help people?

I have a feeling that we need to be real with ourselves and solve problems and not paper over them. I feel like people generally expect search engines to tell them what's really there instead of what people wish were there. And if the engines do that, people can get agitated!

I'd almost say that hurt feelings are prerequisite for real change, hard though that may be.

These are all really interesting questions brought up by this technology, thanks for your thoughts. Disclaimer, I'm a fucking idiot with no idea what I'm talking about.

Re: Imagen, a text-to-image diffusion model

#232
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

Literally the same thing could be said about Google images, but google images is obviously avaliable to the public.

Google knows this will be an unlimited money generator so they're keeping a lid on it.

Re: Imagen, a text-to-image diffusion model

#233

Earlier quoted context omitted.

Translation: we need to hand-tune this to not reflect reality but instead the world as we (Caucasian/Asian male American woke upper-middle class San Fransisco engineers) wish it to be. Maybe that's a nice thing, I wouldn't say their values are wrong but let's call a spade a spade.

"Reality" as defined by the available training set isn't necessarily reality. For example, Google's image search results pre-tweaking had some interesting thoughts on what constitutes a professional hairstyle, and that searches for "men" and "women" should only return light-skinned people: https://www.theguardian.com/technology/2016/apr/08/does-goog... Does that reflect reality? No. (I suspect there are also mostly u…

You know, it wouldn't surprise me if people talking about how black curly hair shouldn't be seen as unprofessional contributed to google thinking there's an association between the concepts of "unprofessional hair" and "black curly hair"

Re: Imagen, a text-to-image diffusion model

#234
post #25

I give it a few years before Google makes stock images irrelevant.

Rolling this into Google Docs seems like a nobrainer.

Google is very conservative about anything that can generate open-ended outputs. Also these models are still very expensive computationally.

Re: Imagen, a text-to-image diffusion model

#235
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

I'm one that welcomes their reasoning. I don't consider myself a social justice kind of guy but I'm not keen on the idea that a tool that is suppose to make life better for everyone has a bias towards one segment of society. This is an important issue(bug?) that needs to be resolved. Specially since there is absolutely no burning reason to release it before it's ready for general use.

Re: Imagen, a text-to-image diffusion model

#236

I know that some monstrous majority of cognitive processing is visual, hence the attention these visually creative models are rightfully getting, but personally I am much more interested in auditory information and would love to see a promptable model for music. Was just listening to "Land Down Under" from Men At Work. Would love to be able to prompt for another artist I have liked: "Tricky playing Land Down Under."…

I believe we’re lacking someone training up a large music model here, but GPT-style transformers can produce music.

gwern can maybe comment here.

An actually scary thing is that AIs are getting okay at reproducing people’s voices.

Re: Imagen, a text-to-image diffusion model

#237

Earlier quoted context omitted.

> If you want the model to understand how the word "nurse" is usually used, without regard for what a "nurse" actually is, then associating it with female is fine. That’s a distinction without a difference. Meaning is use.

Very certainly not, since use is individual and thus a function of competence. So, adherence to meaning depends on the user. Conflict resolution? And anyway - contextually -, the representational natures of "use" (instances) and that of "meaning" (definition) are completely different.

Humans overwhelmingly learn meaning by use, not by definition.

Re: Imagen, a text-to-image diffusion model

#238
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

It’s wild to me that the HN consensus is so often that 1) discourse around the internet is terrible, it’s full of spam and crap, and the internet is an awful unrepresentative snapshot of human existence, and 2) the biases of general-internet-training-data are fine in ML models because it just reflects real life.

It's wild to me that you'd say that. The people complaining (1) aren't following it up with "so we should make sure to restrict the public from internet access entirely". -- that's what would be required to make your juxtaposition make sense.

Moreover, the model doing things like exclusively producing white people when asked to create images of people home brewing beer is "biased" but it's a bias that presumably reflects reality (or at least the internet), if not the reality we'd prefer. Bias means more than "spam and crap", in the ML community bias can also simply mean _accurately_ modeling the underlying distribution when reality falls short of the author's hopes.

For example, if you're interested in learning about what home brewing is the fact that it uses white people would be at least a little unfortunate since there is nothing inherently white and some home brewers aren't white. But if, instead, you wanted to just generate typical home brewing images doing anything but would generate conspicuously unrepresentative images.

But even ignoring the part of the biases which are debatable or of application-specific impact, saying something is unfortunate and saying people should be denied access are entirely different things.

I'll happily delete this comment if you can bring to my attention a single person who has suggested that we lose access to the internet because of spam and crap who has also argued that the release of an internet-biased ML model shouldn't be withheld.

Re: Imagen, a text-to-image diffusion model

#239

I know that some monstrous majority of cognitive processing is visual, hence the attention these visually creative models are rightfully getting, but personally I am much more interested in auditory information and would love to see a promptable model for music. Was just listening to "Land Down Under" from Men At Work. Would love to be able to prompt for another artist I have liked: "Tricky playing Land Down Under."…

I agree. How cool would it be to get an 8 min version of your favorite song? Or an instant DnB remix? Or 10 more songs in the style of your favorite album?

Yeah. I particularly love covers and often can hear in my head X playing Y's song. Would love tools to experiment with that for real.

In practice, my guess is that even though Dall-e level performance in music generation would be stunning and incredible, it would also be tiresome and predictable to consume on any extended basis. I mean- that's my reaction to Dall-e- I find the images astonishing and magical but can only look at them for limited periods of time. At these early stages in this new world the outputs of real individual brains are still more interesting.

But having tools like this to facilitate creation and inspiration by those brains- would be so so cool.

Re: Imagen, a text-to-image diffusion model

#240
As someone who has a layman's understanding of neural networks, and who did some neural network programming ~20 years ago before the real explosion of the field, can someone point to some resources where I can get a better understanding about how this magic works?

I mean, from my perspective, the skill in these (and DALL-E's) image reproductions is truly astonishing. Just looking for more information about how the software actually works, even if there are big chunks of it that are "this is beyond your understanding without taking some in-depth courses".

Post reply on HN