Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

51–60 of 661 posts

Re: Imagen, a text-to-image diffusion model

#51

Is there a way to try this out? DALL-E2 also had amazing demos but the limitations became apparent once real people had a chance to run their own queries.

Looks like no, "The potential risks of misuse raise concerns regarding responsible open-sourcing of code and demos. At this time we have decided not to release code or a public demo. In future work we will explore a framework for responsible externalization that balances the value of external auditing with the risks of unrestricted open-access."

Re: Imagen, a text-to-image diffusion model

#52
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

Translation: we need to hand-tune this to not reflect reality but instead the world as we (Caucasian/Asian male American woke upper-middle class San Fransisco engineers) wish it to be. Maybe that's a nice thing, I wouldn't say their values are wrong but let's call a spade a spade.

Translation: AI has the potential to transform society. When we release this model to the public it will be used in ways we haven’t anticipated. We know the model has bias and we need more time to consider releasing this to the public out of concerns that this transformative technology further perpetuate mistakes that we’ve made in our recent past.

Re: Imagen, a text-to-image diffusion model

#53

Earlier quoted context omitted.

How can we prepare for this? This will result in mass social unrest.

You think so? I'm very high on the Kool-Aid, image generation and text transformation models are core parts of my workflow. (Midjourney, GPT-3) It's still an unruly 7 year old at best. Results need to be verified. Prompt engineering and a sense of creativity are core competencies.

> Prompt engineering and a sense of creativity are core competencies.

It's funny that people are also prompting each other. Parents, friends, teachers, doctors, priests, politicians, managers and marketers are all prompting (advising) us to trigger desired behaviour. Powerful stuff - having a large model and knowing how to prompt it.

Re: Imagen, a text-to-image diffusion model

#54
post #24
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

The big labs have become very sensitive with large model releases. It's too easy to make them generate bad PR, to the point of not releasing almost any of them. Flamingo was also a pretty great vison-language model that wasn't released, not even in a demo. PaLM is supposedly better than GPT-3 but closed off. It will probably take a year for open source models to appear.

The largest models which generate the headline benchmarks are never released after any number of years, it seems.

Very difficult to replicate results.

Re: Imagen, a text-to-image diffusion model

#55
post #41

Really impressive. If we are able to generate such detailed images, is there anything similar for text to music? I would I though that it would be simpler to achieve than text to image.

Compare the size of a raw image file to a raw music file, to get an idea of the complexity difference.

Re: Imagen, a text-to-image diffusion model

#58

Earlier quoted context omitted.

How can we prepare for this? This will result in mass social unrest.

You think so? I'm very high on the Kool-Aid, image generation and text transformation models are core parts of my workflow. (Midjourney, GPT-3) It's still an unruly 7 year old at best. Results need to be verified. Prompt engineering and a sense of creativity are core competencies.

[deleted]

Re: Imagen, a text-to-image diffusion model

#59
post #51

Is there a way to try this out? DALL-E2 also had amazing demos but the limitations became apparent once real people had a chance to run their own queries.

Looks like no, "The potential risks of misuse raise concerns regarding responsible open-sourcing of code and demos. At this time we have decided not to release code or a public demo. In future work we will explore a framework for responsible externalization that balances the value of external auditing with the risks of unrestricted open-access."

> the risks of unrestricted open-access

What exactly is the risk?

Post reply on HN