Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

151–160 of 661 posts

Re: Imagen, a text-to-image diffusion model

#151
post #59
post #51

Earlier quoted context omitted.

Looks like no, "The potential risks of misuse raise concerns regarding responsible open-sourcing of code and demos. At this time we have decided not to release code or a public demo. In future work we will explore a framework for responsible externalization that balances the value of external auditing with the risks of unrestricted open-access."

> the risks of unrestricted open-access What exactly is the risk?

A variation on the axiom "you cannot idiot proof something because there's always a bigger idiot"

Re: Imagen, a text-to-image diffusion model

#152
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

Much like OpenAIs marketing speak about withholding their models for safety, this is just a progressive-sounding cover story for them not wanting to essentially give away a model they spent thousands of man hours and tens of millions of dollars worth of compute training.

Re: Imagen, a text-to-image diffusion model

#153
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

It’s wild to me that the HN consensus is so often that 1) discourse around the internet is terrible, it’s full of spam and crap, and the internet is an awful unrepresentative snapshot of human existence, and 2) the biases of general-internet-training-data are fine in ML models because it just reflects real life.

Re: Imagen, a text-to-image diffusion model

#154
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

Good lord. Withheld? They've published their research, they just aren't making the model available immediately, waiting until they can re-implement it so that you don't get racial slurs popping up when you ask for a cup of "black coffee." >While a subset of our training data was filtered to removed noise and undesirable content, such as pornographic imagery and toxic language, we also utilized LAION-400M dataset whic…

I wonder why they don't like the idea of autogenerated porn... They're already putting most artists out of a job, why not put porn stars out of a job too?

Re: Imagen, a text-to-image diffusion model

#155
I know that some monstrous majority of cognitive processing is visual, hence the attention these visually creative models are rightfully getting, but personally I am much more interested in auditory information and would love to see a promptable model for music. Was just listening to "Land Down Under" from Men At Work. Would love to be able to prompt for another artist I have liked: "Tricky playing Land Down Under." I know of various generative music projects, going back decades, and would appreciate pointers, but as far as I am aware we are still some ways from Imagen/Dalle for music?

Re: Imagen, a text-to-image diffusion model

#156
post #38
post #8

Great. Now even if I do get a Dall-E 2 invite I'll still feel like I'm missing out!

It's always the same with AI research: "we have something amazing but you can't use it because it's too powerful and we think you are an idiot who cannot use your own judgement."

As someone that spent an evening trying to generate images of Hitler Lego I think they have a point.

Re: Imagen, a text-to-image diffusion model

#157
post #85

Earlier quoted context omitted.

If you type as a prompt "most beautiful woman in the world", you get a brown-skinned brown-haired woman with hazel eyes. What should be the right answer then ? You put a blonde, you offend the brown haired. You put blue eyes, you offend the brown eyes. etc.

That's an unanswerable question. Perhaps the answer is "don't". Siri takes this approach for a wide range of queries.

How do you pick what should and shouldn't be restricted? Is there some "offense threshold"? I suspect all queries relating to religion, ethnicity, sexuality, and gender will need to be restricted, which almost certainly means you probably can't include humans at all, other than ones artificially inserted with mathematically proven random attributes. Maybe that's why none are in this demo.

Re: Imagen, a text-to-image diffusion model

#158
post #42

Earlier quoted context omitted.

This raises some really interesting questions. We certainly don't want to perpetuate harmful stereotypes. But is it a flaw that the model encodes the world as it really is, statistically, rather than as we would like it to be? By this I mean that there are more light-skinned people in the west than dark, and there are more women nurses than men, which is reflected in the model's training data. If the model only gener…

It depends on whether you'd like the model to learn casual or correlative relationships. If you want the model to understand what a "nurse" actually is, then it shouldn't be associated with female. If you want the model to understand how the word "nurse" is usually used, without regard for what a "nurse" actually is, then associating it with female is fine. The issue with a correlative model is that it can easily be…

At the end of a day, if you ask for a nurse, should the model output a male or female by default? If the input text lacks context/nuance, then the model must have some bias to infer the user's intent. This holds true for any image it generates; not just the politically sensitive ones. For example, if I ask for a picture of a person, and don't get one with pink hair, is that a shortcoming of the model?

I'd say that bias is only an issue if it's unable to respond to additional nuance in the input text. For example, if I ask for a "male nurse" it should be able to generate the less likely combination. Same with other races, hair colors, etc... Trying to generate a model that's "free of correlative relationships" is impossible because the model would never have the infinitely pedantic input text to describe the exact output image.

Re: Imagen, a text-to-image diffusion model

#159
post #74

Metacalculus, a mass forecasting site, has steadily brought forward the prediction date for a weakly general AI. Jaw-dropping advances like this, only increase my confidence in this prediction. "The future is now, old man." https://www.metaculus.com/questions/3479/date-weakly-general...

I don't see how this gets us (much) closer to general AI. Where is the reasoning?

Perhaps the confluence of NLP and something generative?

Re: Imagen, a text-to-image diffusion model

#160

I give it a few years before Google makes stock images irrelevant.

I really expect them to first make DALL-E and competing networks unfit for commercialization by providing the better choice for free, have stock companies cry in the corner, to just sunset the product a year or two down the road and we're left wandering what to do.
Post reply on HN