Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

291–300 of 661 posts

Re: Imagen, a text-to-image diffusion model

#292
post #25

Earlier quoted context omitted.

Rolling this into Google Docs seems like a nobrainer.

Google is very conservative about anything that can generate open-ended outputs. Also these models are still very expensive computationally.

They're expensive to train, but not awfully expensive to use. Especially if you have hundreds of images you want to generate (due to the way compute devices tend to get much more efficiency with a large batch size).

Google could totally afford it, especially if the feature was hidden behind a button the user had to click, and not just run for every image search.

Re: Imagen, a text-to-image diffusion model

#293

Earlier quoted context omitted.

I wonder why they don't like the idea of autogenerated porn... They're already putting most artists out of a job, why not put porn stars out of a job too?

There's definitely a market for autogenerated porn. But automated porn in a Google branded model for general use around stuff that isn't necessarily intended to be pornographic, on the other hand...

That’s a difficult product because porn is very personalized and if the product is just a little off in latent space it’s going to turn you off.

Also, people have been commenting assuming Google doesn’t want to offend their users or non-users, but they also don’t want to offend their own staff. If you run a porn company you need to hire people okay with that from the start.

Re: Imagen, a text-to-image diffusion model

#294
post #193

Earlier quoted context omitted.

At the end of a day, if you ask for a nurse, should the model output a male or female by default? If the input text lacks context/nuance, then the model must have some bias to infer the user's intent. This holds true for any image it generates; not just the politically sensitive ones. For example, if I ask for a picture of a person, and don't get one with pink hair, is that a shortcoming of the model? I'd say that bi…

This type of bias sounds a lot easier to explain away as a non-issue when we are using "nurse" as the hypothetical prompt. What if the prompt is "criminal", "rapist", or some other negative? Would that change your thought process or would you be okay with the system always returning a person of the same race and gender that statistics indicate is the most likely? Do you see how that could be a problem?

It's an unfortunate reflection of reality. There are three possible outcomes:

1. The model provides a reflection of reality, as politically inconvenient and hurtful as it may be.

2. The model provides an intentionally obfuscated version with either random traits or non correlative traits.

3. The model refuses to answer.

Which of these is ideal to you?

Re: Imagen, a text-to-image diffusion model

#295
post #150

Earlier quoted context omitted.

> If you want the model to understand how the word "nurse" is usually used, without regard for what a "nurse" actually is, then associating it with female is fine. That’s a distinction without a difference. Meaning is use.

Not really; the gender of a nurse is accidental, other properties are essential.

Not really what? How does that contradict what I've said?

Re: Imagen, a text-to-image diffusion model

#297
post #231

Earlier quoted context omitted.

Not the person you responded to, but I do see how someone could be hurt by that, and I want to avoid hurting people. But is this the level at which we should do it? Could skewing search results, i.e. hiding the bias of the real world, give us the impression that everything is fine and we don't need to do anything to actually help people? I have a feeling that we need to be real with ourselves and solve problems and n…

> Could skewing search results, i.e. hiding the bias of the real world Which real world? The population you sample from is going to make a big difference. Do you expect it to reflect your day to day life in your own city? Own country? The entire world? Results will vary significantly.

I'd say it doesn't actually matter, as long as the population sampled is made clear to the user.

If I ask for pictures of Japanese people, I'm not shocked when all the results are of Japanese people. If I asked for "criminals in the United States" and all the results are black people, that should concern me, not because the data set is biased but because the real world is biased and we should do something about that. The difference is that I know what set I'm asking for a sample from, and I can react accordingly.

Re: Imagen, a text-to-image diffusion model

#298

As someone who has a layman's understanding of neural networks, and who did some neural network programming ~20 years ago before the real explosion of the field, can someone point to some resources where I can get a better understanding about how this magic works? I mean, from my perspective, the skill in these (and DALL-E's) image reproductions is truly astonishing. Just looking for more information about how the so…

Figure A.4 in the linked paper is a good high level overview of this model. Shame it was hidden away on page 19 in the appendix! Each box you see there has a section in the paper explaining it in more detail.

Uhh, yeah, I'm going to need much more of an ELI5 than that! Looking at Figure A.4, I understand (again, at a very high-level) the first step of "Frozen Text Encoder", and I have a decent understanding of the upsampling techniques used in the last 2 diffusion model steps, but the middle "Text-to-Image Diffusion Model" step that magically outputs a 64x64 pixel image of an actual golden retriever wearing an actual blue checkered beret and red-dotted turtleneck is where I go "WTF??".

Re: Imagen, a text-to-image diffusion model

#299
post #272
post #223

Earlier quoted context omitted.

Yeah, it seems like it. But it's still just complicated statistical models. Again, where is the reasoning?

I still think we're missing some fundamental insights on how layered planning/forecasting/deducting/reasoning works, and that figuring this out will be necessary in order to create AI that we could say "reasons". But with the recent advances/demonstrations, it seems more likely today than in 2019 that our current computational resources are sufficient to perform magnificantly spooky stuff if they're used correctly. T…

That's a well-balanced response that I can agree with.

I'm an not AGI-skeptic. I'm just a bit skeptical that the topic of this thread is the path forward. It seems to me like an exotic detour.

And, of course intelligence isn't magic. We're producing new intelligent entities at rate of a about ~5 per second globally, every day.

> Does figuring it out seem likely to be many decades away?

1-7?

Re: Imagen, a text-to-image diffusion model

#300

Earlier quoted context omitted.

> If you want the model to understand how the word "nurse" is usually used, without regard for what a "nurse" actually is, then associating it with female is fine. That’s a distinction without a difference. Meaning is use.

Very certainly not, since use is individual and thus a function of competence. So, adherence to meaning depends on the user. Conflict resolution? And anyway - contextually -, the representational natures of "use" (instances) and that of "meaning" (definition) are completely different.

Definition is an entirely artificial construct and doesn't equate to meaning. Definition depends on other words that you also have to understand.
Post reply on HN