Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

31–40 of 661 posts

Re: Imagen, a text-to-image diffusion model

#32
post #3

Jesus Christ. Unlike DALL-E 2, it gets the details right. It also can generate text. The quality is insanely good. This is absolutely mental.

Yes, the posted results are really good, but since we can't play with it we don't know how much cherry picking has been done.

Re: Imagen, a text-to-image diffusion model

#35

The big thing I’m noticing over DALL-E is that it seems to be better at relative positioning. In a MKBHD video about DALLE it would get the elements but not always in the right order. I know google curated some specific images but it seems to be doing a better job there.

Totally—Imagen seems better at composition and relative positioning and text, while DALL-E seems better at lighting, backgrounds, and general artistry.

Re: Imagen, a text-to-image diffusion model

#37
post #29

Earlier quoted context omitted.

The ironic part is that these "social and cultural biases" are purely from a Western, American lens. The people writing that paragraph are completely oblivious to the idea that there could be other cultures other than the Western American one. In attempting to prevent "encoding of social and cultural biases" they have encoded such biases themselves into their own research.

It seems you've got it backwards: "tendency for images portraying different professions to align with Western gender stereotypes" means that they are calling out their own work precisely because it is skewed in the direction of Western American biases.

Yes, the idea is that just because it doesn't align to Western ideals of what seems unbiased doesn't mean that the same is necessarily true for other cultures, and by failing to release the model because it doesn't conform to Western, left wing cultural expectations, the authors are ignoring the diversity of cultures that exist globally.

Re: Imagen, a text-to-image diffusion model

#38
post #8

Great. Now even if I do get a Dall-E 2 invite I'll still feel like I'm missing out!

It's always the same with AI research: "we have something amazing but you can't use it because it's too powerful and we think you are an idiot who cannot use your own judgement."

Re: Imagen, a text-to-image diffusion model

#39
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

Transformers are parallelize-able, right? What’s stopping a large group of people from pooling their compute power together and working towards something like this? IIRC there were some crypto projects a while back that we’re trying to create something similar (golem?)

There are the Eleuther.ai and BigScience projects working on public foundation models. They have a few releases already and currently training GPT-3 sized models.

Re: Imagen, a text-to-image diffusion model

#40
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

Translation: we need to hand-tune this to not reflect reality but instead the world as we (Caucasian/Asian male American woke upper-middle class San Fransisco engineers) wish it to be. Maybe that's a nice thing, I wouldn't say their values are wrong but let's call a spade a spade.

"Reality" as defined by the available training set isn't necessarily reality.

For example, Google's image search results pre-tweaking had some interesting thoughts on what constitutes a professional hairstyle, and that searches for "men" and "women" should only return light-skinned people: https://www.theguardian.com/technology/2016/apr/08/does-goog...

Does that reflect reality? No.

(I suspect there are also mostly unstated but very real concerns about these being used as child pornography, revenge porn, "show my ex brutally murdered" etc. generators.)

Post reply on HN