Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

311–320 of 661 posts

Re: Imagen, a text-to-image diffusion model

#311
post #230

Earlier quoted context omitted.

It’s the same as with an artist: “hey artist, draw me a nurse.” “Hmm okay, do you want it a guy or girl?” “Don’t ask me, just draw what I’m saying.” The artist can then say: “Okay, but accept my biases.” or “I can’t since your input is ambiguous.” For a one-shot generative algorithm you must accept the artist’s biases.

Revert back to average representation of a nurse (give no weight to unspecified criterias, gender, age, skin-color, religion, country, hair-style, no style whether it's a drawing or a photography, no information about the year it was made, etc). “hey artist, draw me a nurse.” “Hmm okay, do you want it a guy or girl?” “Don’t ask me, just draw what I’m saying.” - Ok, I'll draw you what an average nurse looks like. - Wa…

The average nurse has three-halfs of a tit.

Re: Imagen, a text-to-image diffusion model

#312

As someone who has a layman's understanding of neural networks, and who did some neural network programming ~20 years ago before the real explosion of the field, can someone point to some resources where I can get a better understanding about how this magic works? I mean, from my perspective, the skill in these (and DALL-E's) image reproductions is truly astonishing. Just looking for more information about how the so…

> I mean, from my perspective, the skill in these (and DALL-E's) image reproductions is truly astonishing. A basic part of it is that neural networks combine learning and memorizing fluidly inside them, and these networks are really really big, so they can memorize stuff good. So when you see it reproduce a Shiba Inu well, don’t think of it as “the model understands Shiba Inus”. Think of it as making a collage out of…

To be clear, I understand the general techniques about (a) how diffusion models can be used to upsample images and generate more photorealistic (or even "cartoon realistic") results and (b) I understand how they can do basic matching of "someone typed in Shiba Inu, look for images of Shiba Inus".

What I don't understand is how they do the composition. E.g. for "A giant cobra snake on a farm. The snake is made out of corn." I think I could understand how it could reproduce the "A giant cobra snake on a farm" part. What I don't understand is how it accurately pictured "The snake is made out of corn." part, when I'm guessing it has never seen images of snakes made out of corn, and the way it combined "snake" with "made out of corn", in a way that is pretty much how I imagined it would look, is the part I'm baffled by.

Re: Imagen, a text-to-image diffusion model

#313

The big thing I’m noticing over DALL-E is that it seems to be better at relative positioning. In a MKBHD video about DALLE it would get the elements but not always in the right order. I know google curated some specific images but it seems to be doing a better job there.

Totally—Imagen seems better at composition and relative positioning and text, while DALL-E seems better at lighting, backgrounds, and general artistry.

Yeah Dall-e looks amazing, to a mysterious degree even with hints of humour and irony, while imagen images look cheap, one dimensional and quite ugly to be honest.

Still amazing that we're at a point where that's the case, they're both incredible developments.

Re: Imagen, a text-to-image diffusion model

#314

I apologize in advance for the elitist-sounding tone. In my defense the people I’m calling elite I have nothing to do with, I’m certainly not talking about myself. Without a fairly deep grounding in this stuff it’s hard to appreciate how far ahead Brain and DM are. Neither OpenAI nor FAIR ever has the top score on anything unless Google delays publication . And short of FAIR? D2 lacrosse. There are exceptions to such…

In short, it’s all about money.

Re: Imagen, a text-to-image diffusion model

#315
post #224

Earlier quoted context omitted.

No, it's coming from a perspective of moral realism. It's an objective moral truth that racial and ethnic biases are bad. Yet most cultures around the world are racist to at least some degree, and to they extent that the cultures do, they are bad. The argument you're making, paraphrased, is that the idea that biases are bad is itself situated in particular cultural norms. While that is true to some degree, from a mor…

You're confused by the double meaning of the word "bias". Here we mean mathematical biases. For example, a good mathematical model will correctly tell you that people in Japan (geographical term) are more likely to be Japanese (ethnic / racial bias). That's not "objectively morally bad", but instead, it's "correct".

Although what you stated is true, it’s actually a short form of a commonly stated untrue statement “98% of Japan is ethnically Japanese”.

1. that comes from a report from 2006.

2. it’s a misreading, it means “Japanese citizens”, and the government in fact doesn’t track ethnicity at all.

Also, the last time I was in Japan (Jan ‘20) there were literally ten times more immigrants everywhere than my previous trip. Japan is full of immigrants from the rest of Asia these days. They all speak perfect Japanese too.

Re: Imagen, a text-to-image diffusion model

#316

Earlier quoted context omitted.

It’s wild to me that the HN consensus is so often that 1) discourse around the internet is terrible, it’s full of spam and crap, and the internet is an awful unrepresentative snapshot of human existence, and 2) the biases of general-internet-training-data are fine in ML models because it just reflects real life.

Why is it wild? How is it contradictory?

If these models spit out the data they were trained on and the training data isn’t representative of reality, then they won’t spit out content that’s representative of reality either.

So people shouldn’t say ‘these concerns are just woke people doing dumb woke stuff, but the model is just reflecting reality.’

Re: Imagen, a text-to-image diffusion model

#317

Earlier quoted context omitted.

Yes actually, subconscious bias due to historical prejudice does have a large effect on society. Obviously there are things with much larger effects, that doesn't mean that this doesn't exist. > Oh no I asked the model to draw a doctor and it drew a male doctor, I guess there's no point in me pursuing medical studies If you don't think this is a real thing that happens to children you're not thinking especially hard.…

> If you don't think this is a real thing that happens to children you're not thinking especially hard I believe that's where parenting comes in. Maybe I'm too cynical but I think that the parents' job is to undo all of the harm done by society and instill in their children the "correct" values.

Isn't that putting an undue load on parents?

It seems extremely unfair that parents of young black men should have to work extra hard to tell their kids they're not destined to be criminals. Hell, it's not fair on parents of blonde girls to tell their kids they don't have to be just dumb and pretty.

(note: I am deliberately picking bad stereotypes that are pervasive in our culture... I am not in any way suggesting those are true.)

Re: Imagen, a text-to-image diffusion model

#318
post #193

Earlier quoted context omitted.

This type of bias sounds a lot easier to explain away as a non-issue when we are using "nurse" as the hypothetical prompt. What if the prompt is "criminal", "rapist", or some other negative? Would that change your thought process or would you be okay with the system always returning a person of the same race and gender that statistics indicate is the most likely? Do you see how that could be a problem?

It's an unfortunate reflection of reality. There are three possible outcomes: 1. The model provides a reflection of reality, as politically inconvenient and hurtful as it may be. 2. The model provides an intentionally obfuscated version with either random traits or non correlative traits. 3. The model refuses to answer. Which of these is ideal to you?

What makes you think those are the only options? Why can't we have an option that the model returns a range of different outputs based off a prompt?

A model that returns 100% of nurses as female might be statistically more accurate than a model that returns 50% of nurses as female, but it is still not an accurate reflection of the real world. I agree that the model shouldn't return a male nurse 50% of the time. Yet an accurate model needs to be able to occasionally return a male nurse without being directly prompted for a "male nurse". Anything else would also be inaccurate.

Re: Imagen, a text-to-image diffusion model

#319

I apologize in advance for the elitist-sounding tone. In my defense the people I’m calling elite I have nothing to do with, I’m certainly not talking about myself. Without a fairly deep grounding in this stuff it’s hard to appreciate how far ahead Brain and DM are. Neither OpenAI nor FAIR ever has the top score on anything unless Google delays publication . And short of FAIR? D2 lacrosse. There are exceptions to such…

> But the incentives are weird here: who, exactly, does it serve for us plebs to be able to train these things from scratch?

I'm not sure it matters. The history of computing shows that within the decade we will all have the ability to train and use these models.

Re: Imagen, a text-to-image diffusion model

#320
post #309

I'm curious why all of these tools seem to be almost tailored toward making meme images? The kind of early 2010's, over the top description of something that's ridiculous

These things can make any image you can define in terms of a corpus of other images. That was true at lower resolution five years ago.

To the extent that they get used for making bored ape images or whatever meme du juor, it says much more about the kind of pictures people want to see.

I personally find the weird deep dreaming dogs with spikes coming out of their heads more mathematically interesting, but I can understand why that doesn’t sell as well.

Post reply on HN