Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

471–480 of 661 posts

Re: Imagen, a text-to-image diffusion model

#471
post #437

I have to wonder how much releasing these models will "poison the well" and fill the internet with AI generated images that make training an improved model difficult. After all if every 9/10 "oil painted" image online starts being from these generative models it'll become increasingly difficult to scrape the web and to learn from real world data in a variety of domains. Essentially once these things are widely availa…

Just yesterday I was speculating that current AI is bad at math because math on the internet is spectacularly terrible. I think you’re right, and it’s unlikely that we (society) will convince people to label their AI content as such so that scraping is still feasible. It’s far more likely that companies will be formed to provide “pristine training sets of human-created content”, and quite likely they will be subscrip…

>“pristine training sets of human-created content”

well, we do have organic/farmed/handcrafted/etc. food. One can imagine information nutrition label - "contains 70% AI generated content, triggers 25% of the daily dopamine release target".

Re: Imagen, a text-to-image diffusion model

#472
post #437

I have to wonder how much releasing these models will "poison the well" and fill the internet with AI generated images that make training an improved model difficult. After all if every 9/10 "oil painted" image online starts being from these generative models it'll become increasingly difficult to scrape the web and to learn from real world data in a variety of domains. Essentially once these things are widely availa…

How would that really happen? It seems to me you're assuming that there's no such thing as extant databases of actual oil paintings, that people will stop producing, documenting, and curating said paintings. I think the internet and curated image databases are far more well kept than your proposed model accounts for.

My hypothetical example is not really about oil paintings, but the fact these models will surely get deployed and used for stock photos for articles, on art pages etc.

I think this will introduce unavoidable background noise that will be super hard to fully eliminate in future large scale data sets scraped from the web, there's always going to be more and more photorealistic pictures of "cats" "chairs" etc. in the data that are close to looking real but not quite, and we can never really go back to a world where there's only "real" pictures, or "authentic human art" on the internet.

Re: Imagen, a text-to-image diffusion model

#473
One thing that no one predicted in AI development was how good it would become at some completely unexpected tasks while being not so great at the ones we supposed/hoped it would be good.

AI was expected to grow like a child. Somehow blurting out things that would show some increasing understanding on a deep level but poor syntax.

In fact we get the exact opposite. AI is creating texts that are syntaxically correct and very decently articulated and pictures that are insanely good.

And these texts and images are created from a text prompt?! There is no way to interface with the model other than by freeform text. That is so weird to me.

Yet it doesn’t feel intelligent at all at first. You can’t ask it to draw “a chess game with a puzzle where white mates in 4 moves”.

Yet sometimes GPT makes very surprising inferences. And it starts to feel like there is something going on a deeper level.

DeepMind’s AlphaXxx models are more in line with how I expected things to go. Software that gets good at expert tasks that we as humans are too limited to handle.

Where it’s headed, we don’t know. But I bet it’s going to be difficult to tell the “intelligence” from the “varnish”

Re: Imagen, a text-to-image diffusion model

#474
post #326
post #132

Interesting and cool technology - but I can't seem to ignore that every high-quality AI art application is always closed, and I don't seem to buy the ethics excuse for that. The same was said for GPT, yet I see nothing but creativity coming out from its users nowadays.

I don't buy the ethics but I do buy the obvious PR nightmare that would inevitably happen if journalists could play with this and immediately publish their findings of "racist imagery generated by googles AI". That's all it's about and us complaining is not going to make them change their minds.

Then they should be honest about it. They can use all the PR lingo they want but don't flat out lie about it.

Lying about ethics or misattributing their actions to some misguided sense of "social" responsibility puts google in a far worse light in my eyes. I can't help but wonder how many skilled employees were driven off from accepting a position at google because of lies like these.

Re: Imagen, a text-to-image diffusion model

#476

Earlier quoted context omitted.

But why is it a problem? The AI is just a mirror showing us ourselves. That’s a good thing. How does it help anyone to make an AI that presents a fake world so that we can pretend that we live in a world that we actually don’t? Disassociation from reality is more dangerous than bias.

In the days when Sussman was a novice Minsky once came to him as he sat hacking at the PDP-6. "What are you doing?", asked Minsky. "I am training a randomly wired neural net to play Tic-Tac-Toe." "Why is the net wired randomly?", asked Minsky. "I do not want it to have any preconceptions of how to play" Minsky shut his eyes, "Why do you close your eyes?", Sussman asked his teacher. "So that the room will be empty." A…

The model makes inferences about the world from training data. When it sees more female nurses than male nurses in its training set, if infers that most nurses are female. This is a correct inference.

If they were to weight the training data so that there were an equal number of male and female nurses, then it may well produce male and female nurses with equal probability, but it would also learn an incorrect understanding of the world.

That is quite distinct from weighting the data so that it has a greater correspondence to reality. For example, if Africa is not represented well then weighting training data from Africa more strongly is justifiable.

The point is, it’s not a good thing for us to intentionally teach AIs a world that is idealized and false.

As these AIs work their way into our lives it is essential that they reproduce the world in all of its grit and imperfections, lest we start to disassociate from reality.

Chinese media (or insert your favorite unfree regime) also presents China as a utopia.

Re: Imagen, a text-to-image diffusion model

#477
post #332

Earlier quoted context omitted.

So, the model should have a knowledge of political correctness, and return multiple results if the first choice might reinforce a stereotype?

I never said anything about political correctness. You implied that you want a model that "provides a reflection of reality". All nurses being female is not "a reflection of reality". It is a distortion of reality because the model doesn't actually understand gender or nurses.

A majority of nurses are women, therefore a woman would be a reasonable representation of a nurse. Obviously that's not a helpful stereotype, because male nurses exist and face challenges due to not fitting the stereotypes. The model is dumb, and outputs what it's seen. Is that wrong?

Re: Imagen, a text-to-image diffusion model

#478
post #437

I have to wonder how much releasing these models will "poison the well" and fill the internet with AI generated images that make training an improved model difficult. After all if every 9/10 "oil painted" image online starts being from these generative models it'll become increasingly difficult to scrape the web and to learn from real world data in a variety of domains. Essentially once these things are widely availa…

I wonder if google images could just seed in some generated images when none relevant are found..

Re: Imagen, a text-to-image diffusion model

#479
post #326
post #132

Interesting and cool technology - but I can't seem to ignore that every high-quality AI art application is always closed, and I don't seem to buy the ethics excuse for that. The same was said for GPT, yet I see nothing but creativity coming out from its users nowadays.

I don't buy the ethics but I do buy the obvious PR nightmare that would inevitably happen if journalists could play with this and immediately publish their findings of "racist imagery generated by googles AI". That's all it's about and us complaining is not going to make them change their minds.

There already is an article calling DALL E racist and it isn’t even public yet. Just imagine what horrible things the general public will get it to spit out.

Re: Imagen, a text-to-image diffusion model

#480
post #43

Earlier quoted context omitted.

Translation: we need to hand-tune this to not reflect reality but instead the world as we (Caucasian/Asian male American woke upper-middle class San Fransisco engineers) wish it to be. Maybe that's a nice thing, I wouldn't say their values are wrong but let's call a spade a spade.

I know you're anon trolling, but the authors' names are: Chitwan Saharia, William Chan, Saurabh Saxena†, Lala Li†, Jay Whang†, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho†, David Fleet†, Mohammad Norouzi

Google AI researchers don't have the final say in what gets published and what doesn't. I think there was a huge controversy when people learned about it last year.
Post reply on HN