Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

11–20 of 661 posts

Re: Imagen, a text-to-image diffusion model

#11
The big thing I’m noticing over DALL-E is that it seems to be better at relative positioning. In a MKBHD video about DALLE it would get the elements but not always in the right order. I know google curated some specific images but it seems to be doing a better job there.

Re: Imagen, a text-to-image diffusion model

#12
Interesting discovery they made

> We show that scaling the pretrained text encoder size is more important than scaling the diffusion model size.

There seems to be an unexpected level of synergy between text and vision models. Can't wait to see what video and audio modalities will add to the mix.

Re: Imagen, a text-to-image diffusion model

#13

Metacalculus, a mass forecasting site, has steadily brought forward the prediction date for a weakly general AI. Jaw-dropping advances like this, only increase my confidence in this prediction. "The future is now, old man." https://www.metaculus.com/questions/3479/date-weakly-general...

How can we prepare for this? This will result in mass social unrest.

You think so? I'm very high on the Kool-Aid, image generation and text transformation models are core parts of my workflow. (Midjourney, GPT-3)

It's still an unruly 7 year old at best. Results need to be verified. Prompt engineering and a sense of creativity are core competencies.

Re: Imagen, a text-to-image diffusion model

#14
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

The ironic part is that these "social and cultural biases" are purely from a Western, American lens. The people writing that paragraph are completely oblivious to the idea that there could be other cultures other than the Western American one. In attempting to prevent "encoding of social and cultural biases" they have encoded such biases themselves into their own research.

Re: Imagen, a text-to-image diffusion model

#17
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

The ironic part is that these "social and cultural biases" are purely from a Western, American lens. The people writing that paragraph are completely oblivious to the idea that there could be other cultures other than the Western American one. In attempting to prevent "encoding of social and cultural biases" they have encoded such biases themselves into their own research.

What makes you think the authors are all American?

Re: Imagen, a text-to-image diffusion model

#19
Does it do partial image reconstruction like DALL-E2? Where you cut out part of an existing image and the neural network can fill it back in.

I believe this type of content generation will be the next big thing or at least one of them. But people will want some customization to make their pictures “unique” and fix AI’s lack of creativity and other various shortcomings. Plus edit out the remaining lapses in logic/object separation (which there are some even in the given examples).

Still, being able to create arbitrary stock photos is really useful and i bet these will flood small / low-budget projects

Re: Imagen, a text-to-image diffusion model

#20
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

The ironic part is that these "social and cultural biases" are purely from a Western, American lens. The people writing that paragraph are completely oblivious to the idea that there could be other cultures other than the Western American one. In attempting to prevent "encoding of social and cultural biases" they have encoded such biases themselves into their own research.

https://en.wikipedia.org/wiki/Moral_relativism
Post reply on HN