Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

541–550 of 661 posts

Re: Imagen, a text-to-image diffusion model

#541
post #297

Earlier quoted context omitted.

> Could skewing search results, i.e. hiding the bias of the real world Which real world? The population you sample from is going to make a big difference. Do you expect it to reflect your day to day life in your own city? Own country? The entire world? Results will vary significantly.

I'd say it doesn't actually matter, as long as the population sampled is made clear to the user. If I ask for pictures of Japanese people, I'm not shocked when all the results are of Japanese people. If I asked for "criminals in the United States" and all the results are black people, that should concern me, not because the data set is biased but because the real world is biased and we should do something about that.…

In a way, if the model brings back an image for "criminals in the United States" that isn't based on the statistical reality, isn't it essentially complicit in sweeping a major social issue under the rug?

We may not like what it shows us, but blindfolding ourselves is not the solution to that problem.

Re: Imagen, a text-to-image diffusion model

#542
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

Literally the same thing could be said about Google images, but google images is obviously avaliable to the public. Google knows this will be an unlimited money generator so they're keeping a lid on it.

Given that there's already many competing models in this space prior to any of them having been brought to market, it seems more likely that it will be commoditized.

Re: Imagen, a text-to-image diffusion model

#544
post #473

One thing that no one predicted in AI development was how good it would become at some completely unexpected tasks while being not so great at the ones we supposed/hoped it would be good. AI was expected to grow like a child. Somehow blurting out things that would show some increasing understanding on a deep level but poor syntax. In fact we get the exact opposite. AI is creating texts that are syntaxically correct a…

I doubt 99% of humans can draw a ”chess game with a puzzle where white mates in 4 moves”

It's surprisingly easy to construct such a thing.

If you want to be trivial about it, you can just have white back-rank mate with a rook, and black has 4 pieces to block with.

Re: Imagen, a text-to-image diffusion model

#546

Earlier quoted context omitted.

Look at carpentry blogs, recipe blogs. Nearly all of it is junk content. I bet if you combined GPT and imagen or dalle2 you could replace all of them. Just provide a betty crocker recipe and let it generate a blog that has weekly updates and even a bunch of images - "happy family enjoying pancakes together" I can see the future as being devoid of any humanity.

I wrote a comedic "Best Apache Chef recipe" article[1] mocking these sites. I guess the concern would be: If one of these recipe websites _was_ generated by an AI, the ingredients _look_ correct to an AI but are otherwise wrong - then what do you do? Baking soda swapped with baking powder. Tablespoons instead of teaspoons. Add 2tbsp of flower to the caramel macchiato. Whoops! Meant sugar. [0] http://slimsag.com/best-…

> Allow server to cool down for ~10 minutes

Epic

Re: Imagen, a text-to-image diffusion model

#547
post #493

Earlier quoted context omitted.

People training newer models just have to look for the "Imagen" tag or the Dall-E2 rainbow at the corner and heuristically exclude images having these. This is trivial. Unless you assume there are bad actors who will crop out the tags. Not many people now have access to Dall-E2 or will have access to Imagen. As someone working in Vision, I am also thinking about whether to include such images deliberately. Using imag…

Most images you see from these services will not have a watermark on them. Cropping is trivial.

Perhaps a watermark should be embedded in a subtle way across the whole image. What is the word? "Steganography" is designed to solve a different problem and I don't think it survives recompression etc. Is there a way to create weakly secure watermarks that are invisible to naked eye, spread across the whole image, resistant to scaling and lossless compression (to a point)?

Re: Imagen, a text-to-image diffusion model

#549

Earlier quoted context omitted.

Look at carpentry blogs, recipe blogs. Nearly all of it is junk content. I bet if you combined GPT and imagen or dalle2 you could replace all of them. Just provide a betty crocker recipe and let it generate a blog that has weekly updates and even a bunch of images - "happy family enjoying pancakes together" I can see the future as being devoid of any humanity.

I wrote a comedic "Best Apache Chef recipe" article[1] mocking these sites. I guess the concern would be: If one of these recipe websites _was_ generated by an AI, the ingredients _look_ correct to an AI but are otherwise wrong - then what do you do? Baking soda swapped with baking powder. Tablespoons instead of teaspoons. Add 2tbsp of flower to the caramel macchiato. Whoops! Meant sugar. [0] http://slimsag.com/best-…

[deleted]

Re: Imagen, a text-to-image diffusion model

#550

Earlier quoted context omitted.

Looking at these… I can’t help but wonder if these are literal examples of AI imagination?

I've started to ask myself if my own creativity is a result of random sampling from the diffusion tapestry of associated memories and experience on that topic.

I do wonder what Dall-E 2 would output for a request along the lines of "A still life of a vase of flowers in a completely new art style."
Post reply on HN