Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

521–530 of 661 posts

Re: Imagen, a text-to-image diffusion model

#521
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

Even as a pretty left leaning person, I gotta agree. We should see AI’s pollution by human shortcoming akin to the fact that our world is the product of many immoralities that came before us. It sucks that they ever existed, but we should understand that the results are, by definition, a product of the past, and let them live in that context.

Re: Imagen, a text-to-image diffusion model

#523
post #493
post #437

I have to wonder how much releasing these models will "poison the well" and fill the internet with AI generated images that make training an improved model difficult. After all if every 9/10 "oil painted" image online starts being from these generative models it'll become increasingly difficult to scrape the web and to learn from real world data in a variety of domains. Essentially once these things are widely availa…

People training newer models just have to look for the "Imagen" tag or the Dall-E2 rainbow at the corner and heuristically exclude images having these. This is trivial. Unless you assume there are bad actors who will crop out the tags. Not many people now have access to Dall-E2 or will have access to Imagen. As someone working in Vision, I am also thinking about whether to include such images deliberately. Using imag…

> Unless you assume there are bad actors who will crop out the tags.

I don't know why people do that but lots of randoms on the internet do that and they're not even bad actors per se. The removed signatures from art posted online became a kind of a meme itself. Especially when comic strips are reposted on Reddit. So yeah, we'll see lots of them.

Re: Imagen, a text-to-image diffusion model

#524
post #370

Earlier quoted context omitted.

Imagen takes text embeddings, OpenAI model takes image embeddings instead, this is the reason. There are other models that can generate text: latent diffusion trained on LAION-400M, GLIDE, DALL-E (1).

My understanding of the terms text and image embeddings is that they are ways of representing text or images as vectors. But, I don't understand how that would help with the process of actually drawing the symbols for those letters.

If the model takes text embeddings/tokens as an input, it can create a connection between the caption and the text on the image (sometimes they are really similar).

Re: Imagen, a text-to-image diffusion model

#525
post #437

I have to wonder how much releasing these models will "poison the well" and fill the internet with AI generated images that make training an improved model difficult. After all if every 9/10 "oil painted" image online starts being from these generative models it'll become increasingly difficult to scrape the web and to learn from real world data in a variety of domains. Essentially once these things are widely availa…

I can see a world where in person consumption of creative media (art, music, movies etc), where all devices are to be left at the door, becomes more and more sought after and lucrative.

If the AI models can't consume it, it can't be commoditised and, well, ruined.

Re: Imagen, a text-to-image diffusion model

#526

Earlier quoted context omitted.

Looking at these… I can’t help but wonder if these are literal examples of AI imagination?

I've started to ask myself if my own creativity is a result of random sampling from the diffusion tapestry of associated memories and experience on that topic.

What else could creativity possible be?

Re: Imagen, a text-to-image diffusion model

#527
post #24

Earlier quoted context omitted.

The big labs have become very sensitive with large model releases. It's too easy to make them generate bad PR, to the point of not releasing almost any of them. Flamingo was also a pretty great vison-language model that wasn't released, not even in a demo. PaLM is supposedly better than GPT-3 but closed off. It will probably take a year for open source models to appear.

That's because we're still bad about long-tailed data and that people outside the research don't realize that we're first prioritizing realistic images before we deal with long-tailed data (which is going to be the more generic form of bias). To be honest, it is a bit silly to focus on long-tailed data when results aren't great. That's why we see the constant pattern of getting good on a dataset and then focusing on…

Well, if you showed that pixelated image to someone who has never seen Obama - would they make him white? I think so.

Re: Imagen, a text-to-image diffusion model

#528

This looks incredible but I do notice that all the images are of a similar theme. Specifically there are no human figures.

I believe DALLE and likely this model excluded images of people so it could not be misused

Interesting, I had not understood that!

Re: Imagen, a text-to-image diffusion model

#530

Earlier quoted context omitted.

> Copenhagen ethics (used by most people) The idea that most people use any coherent ethical framework (even something as high level and nearly content-free as Copenhagen) much less a particular coherent ethical framework is, well, not well supported by the evidence. > require that all negative outcomes of a thing X become yours if you interact with X. It is not sensible to interact with high negativity things unless…

I'm sure you are capable of steelmanning the argument.

Maybe: Most peoples morals require that all negative outcomes of a thing X become yours if you interact with X.

I am not sure of the evidence but that would seem almost right.

Except for, for example a story I read where a couple lost their housing deposit due to a payment timing issue. They used a lawyer and were not doing anything “fancy” like buying via a holding company. They interacted with “buying a house”, so is this just tough shit because they interacted with X.

That sounds like the original Bitcoin “not your keys not your coin” kind of morality.

I don’t think I can figure out the steel man.

Post reply on HN