Imagen, a text-to-image diffusion model
31–40 of 661 posts
Re: Imagen, a text-to-image diffusion model
#32Jesus Christ. Unlike DALL-E 2, it gets the details right. It also can generate text. The quality is insanely good. This is absolutely mental.
Re: Imagen, a text-to-image diffusion model
#33Re: Imagen, a text-to-image diffusion model
#34OpenAi really thought they had done something with DALL-E, then Google's all "hold my beer".
Re: Imagen, a text-to-image diffusion model
#35The big thing I’m noticing over DALL-E is that it seems to be better at relative positioning. In a MKBHD video about DALLE it would get the elements but not always in the right order. I know google curated some specific images but it seems to be doing a better job there.
Re: Imagen, a text-to-image diffusion model
#36Oh well.
Re: Imagen, a text-to-image diffusion model
#37Earlier quoted context omitted.
The ironic part is that these "social and cultural biases" are purely from a Western, American lens. The people writing that paragraph are completely oblivious to the idea that there could be other cultures other than the Western American one. In attempting to prevent "encoding of social and cultural biases" they have encoded such biases themselves into their own research.
It seems you've got it backwards: "tendency for images portraying different professions to align with Western gender stereotypes" means that they are calling out their own work precisely because it is skewed in the direction of Western American biases.
Re: Imagen, a text-to-image diffusion model
#38Great. Now even if I do get a Dall-E 2 invite I'll still feel like I'm missing out!
Re: Imagen, a text-to-image diffusion model
#39>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…
Transformers are parallelize-able, right? What’s stopping a large group of people from pooling their compute power together and working towards something like this? IIRC there were some crypto projects a while back that we’re trying to create something similar (golem?)
Re: Imagen, a text-to-image diffusion model
#40>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…
Translation: we need to hand-tune this to not reflect reality but instead the world as we (Caucasian/Asian male American woke upper-middle class San Fransisco engineers) wish it to be. Maybe that's a nice thing, I wouldn't say their values are wrong but let's call a spade a spade.
For example, Google's image search results pre-tweaking had some interesting thoughts on what constitutes a professional hairstyle, and that searches for "men" and "women" should only return light-skinned people: https://www.theguardian.com/technology/2016/apr/08/does-goog...
Does that reflect reality? No.
(I suspect there are also mostly unstated but very real concerns about these being used as child pornography, revenge porn, "show my ex brutally murdered" etc. generators.)