Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

581–590 of 661 posts

Re: Imagen, a text-to-image diffusion model

#581

Earlier quoted context omitted.

Is this a joke?

No

To expand a bit for the grandparent, if you check out this authors other repos you'll notice they have a thing for implementing these papers (multiple DALLE-2 implementations for instance). You should expect to see an implementation there pretty quickly I'd guess.

Re: Imagen, a text-to-image diffusion model

#582
post #560

Can anybody give me short high-level explanation how the model achieves these results? I'm especially interested in the image synthesis, not the language parsing. For example, what kind of source images are used for the snake made of corn[0]? It's baffling to me how the corn is mapped to the snake body. [0] https://gweb-research-imagen.appspot.com/main_gallery_images...

In the paper they say about half the training data was an internal training set, and the other half came from: https://laion.ai/laion-400-open-dataset/

Re: Imagen, a text-to-image diffusion model

#583
post #522

I find it a bit disturbing that they talk about social impact of totally imaginary pictures of racoon. Of course, working in a golden lab at Google may twist your views on society.

Oh, I would say they are probably underestimating the impact. You only saw the images they thought couldn't raise alarm bells. Anyone will be able to create photorealistic images of anyone doing anything, Anything! This is certainly a dangerous and society altering tech. It won't all be teddy bears and racoons playing poker.

Re: Imagen, a text-to-image diffusion model

#584
post #437

I have to wonder how much releasing these models will "poison the well" and fill the internet with AI generated images that make training an improved model difficult. After all if every 9/10 "oil painted" image online starts being from these generative models it'll become increasingly difficult to scrape the web and to learn from real world data in a variety of domains. Essentially once these things are widely availa…

Look at carpentry blogs, recipe blogs. Nearly all of it is junk content. I bet if you combined GPT and imagen or dalle2 you could replace all of them. Just provide a betty crocker recipe and let it generate a blog that has weekly updates and even a bunch of images - "happy family enjoying pancakes together" I can see the future as being devoid of any humanity.

I see the opposite future.

As AI advances, a lot of people will look after experiencing life outside the digital world.

Even digital communication will not be trustworthy anymore with deepfaces and everything else, so people will want to get together more often.

Edit: for the lazy ones, yeah, digital will be a sad and heartless environment...

Re: Imagen, a text-to-image diffusion model

#585
post #463

Earlier quoted context omitted.

The irony is that when the majority of content becomes computer-generated, most of that content will also be computer-consumed. Neil Stephenson covered this briefly in "Fall; or Dodge In Hell." So much 'net content was garbage, AI-generated, and/or spam that it could only be consumed via "editors" (either AI or AI+human, depending on your income level) that separated the interesting sliver of content from...everythin…

He was definitely onto something in that book where people also resort to using blockchains to fingerprint their behavior and build an unbreakable chain of authenticity. Later in that book that is used to authorize the hardware access of the deceased and uploaded individuals. A bit far out there in terms of plot but the notion of authenticating based on a multitude of factors and fingerprints is not that strange. We'…

> blockchains to fingerprint their behavior and build an unbreakable chain of authenticity. Later in that book that is used to authorize the hardware access of the deceased and uploaded individuals.

maybe I misunderstood, but I had it that people used generative AI models that would transform the media they produced. The generated content can be uniquely identified, but the creator (or creators) retains anonymity. Later these generative AI models morphed into a form of identity since they could be accurately and uniquely identified.

Re: Imagen, a text-to-image diffusion model

#586
post #437

I have to wonder how much releasing these models will "poison the well" and fill the internet with AI generated images that make training an improved model difficult. After all if every 9/10 "oil painted" image online starts being from these generative models it'll become increasingly difficult to scrape the web and to learn from real world data in a variety of domains. Essentially once these things are widely availa…

[deleted]

Re: Imagen, a text-to-image diffusion model

#587

For people complaining that they can't play with the model... I work at Google and I also can't play with the model :'(

I think they address some of the reasoning behind this pretty clearly in the write-up as well?

> The potential risks of misuse raise concerns regarding responsible open-sourcing of code and demos. At this time we have decided not to release code or a public demo. In future work we will explore a framework for responsible externalization that balances the value of external auditing with the risks of unrestricted open-access.

I can see the argument here. It would be super fun to test this model's ability to generate arbitrary images, but "arbitrary" also contains space for a lot of distasteful stuff. Add in this point:

> While a subset of our training data was filtered to removed noise and undesirable content, such as pornographic imagery and toxic language, we also utilized LAION-400M dataset which is known to contain a wide range of inappropriate content including pornographic imagery, racist slurs, and harmful social stereotypes. Imagen relies on text encoders trained on uncurated web-scale data, and thus inherits the social biases and limitations of large language models. As such, there is a risk that Imagen has encoded harmful stereotypes and representations, which guides our decision to not release Imagen for public use without further safeguards in place.

That said, I hope they're serious about the "framework for responsible externalization" part, both because it would be really fun to play with this model and because it would be interesting to test it outside of their hand-picked examples.

Re: Imagen, a text-to-image diffusion model

#588

Earlier quoted context omitted.

Look at carpentry blogs, recipe blogs. Nearly all of it is junk content. I bet if you combined GPT and imagen or dalle2 you could replace all of them. Just provide a betty crocker recipe and let it generate a blog that has weekly updates and even a bunch of images - "happy family enjoying pancakes together" I can see the future as being devoid of any humanity.

I see the opposite future. As AI advances, a lot of people will look after experiencing life outside the digital world. Even digital communication will not be trustworthy anymore with deepfaces and everything else, so people will want to get together more often. Edit: for the lazy ones, yeah, digital will be a sad and heartless environment...

This is my theory as well. There'll be a short period where some of us at the forefront will enrich themselves by flooding the internet with imagery never seen before. That'll be a bubble where people think "abundance" has been solved but then it'll pop as people start to not trust anything they see online anymore and as you say, only trust and interact with things in the real world (wouldn't surpise me if regulation got involved here too somehow).

Re: Imagen, a text-to-image diffusion model

#589
post #506

Earlier quoted context omitted.

Look at carpentry blogs, recipe blogs. Nearly all of it is junk content. I bet if you combined GPT and imagen or dalle2 you could replace all of them. Just provide a betty crocker recipe and let it generate a blog that has weekly updates and even a bunch of images - "happy family enjoying pancakes together" I can see the future as being devoid of any humanity.

Doesn't it increases the value of genuine human-produced content? Or their NFTs!

For the very skilled yes. But a lot of low skilled artists of content creators will have the rugs pulled out from under them. (And how will we ever get high skilled artists trained in the future if they can't make a living from their lower tier output before they reach mastery.)

Re: Imagen, a text-to-image diffusion model

#590

Earlier quoted context omitted.

In the days when Sussman was a novice Minsky once came to him as he sat hacking at the PDP-6. "What are you doing?", asked Minsky. "I am training a randomly wired neural net to play Tic-Tac-Toe." "Why is the net wired randomly?", asked Minsky. "I do not want it to have any preconceptions of how to play" Minsky shut his eyes, "Why do you close your eyes?", Sussman asked his teacher. "So that the room will be empty." A…

The model makes inferences about the world from training data. When it sees more female nurses than male nurses in its training set, if infers that most nurses are female. This is a correct inference. If they were to weight the training data so that there were an equal number of male and female nurses, then it may well produce male and female nurses with equal probability, but it would also learn an incorrect underst…

> As these AIs work their way into our lives it is essential that they reproduce the world in all of its grit and imperfections...

Is it? I'm reminded of the Microsoft Tay experiment, were they attempted to train an AI by letting Twitter users interact with it.

The result was a non-viable mess that nobody liked.

Post reply on HN