I have to wonder how much releasing these models will "poison the well" and fill the internet with AI generated images that make training an improved model difficult. After all if every 9/10 "oil painted" image online starts being from these generative models it'll become increasingly difficult to scrape the web and to learn from real world data in a variety of domains. Essentially once these things are widely availa…
Imagen, a text-to-image diffusion model
641–650 of 661 posts
Re: Imagen, a text-to-image diffusion model
#642Earlier quoted context omitted.
I think they address some of the reasoning behind this pretty clearly in the write-up as well? > The potential risks of misuse raise concerns regarding responsible open-sourcing of code and demos. At this time we have decided not to release code or a public demo. In future work we will explore a framework for responsible externalization that balances the value of external auditing with the risks of unrestricted open-…
I find this is a really bad precedent. Not far from now they'll achieve "super human" level general AI, and be like "yeah it's too powerful for you, we'll keep it internal".
Re: Imagen, a text-to-image diffusion model
#643Re: Imagen, a text-to-image diffusion model
#644Earlier quoted context omitted.
At the end of a day, if you ask for a nurse, should the model output a male or female by default? If the input text lacks context/nuance, then the model must have some bias to infer the user's intent. This holds true for any image it generates; not just the politically sensitive ones. For example, if I ask for a picture of a person, and don't get one with pink hair, is that a shortcoming of the model? I'd say that bi…
> At the end of a day, if you ask for a nurse, should the model output a male or female by default? Randomly pick one. > Trying to generate a model that's "free of correlative relationships" is impossible because the model would never have the infinitely pedantic input text to describe the exact output image. Sure, and you can never make a medical procedure 100% safe. Doesn't mean that you don't try to make them safe…
I say this because I’ve been visiting a number of childcare centres over the past few days and I still have yet to see a single male teacher.
Re: Imagen, a text-to-image diffusion model
#645Earlier quoted context omitted.
Compare the size of a raw image file to a raw music file, to get an idea of the complexity difference.
Think sheet music, not an mp3
Re: Imagen, a text-to-image diffusion model
#646Earlier quoted context omitted.
> the risks of unrestricted open-access What exactly is the risk?
"Make a photograph of Joe Biden in a hotel room bed with Kim Jong-un." Simply the ease at which people are going to be able to make extremely-realistic game photographs is going to do some damage to the world. It's inevitable, but it might be good to postpone it.
I don't understand why. If someone has gone to a blockbuster movie in the last 15 years, they're very familiar with the concept of making people, sets, and entire worlds, that don't exist, with photorealistic accuracy. Being able to make fictitious photorealistic images isn't remotely a new ability, it's just an ability that's now automated.
If this is released, I think any damage would be extremely fleeting, as people pumped out thousands of these images, and people grow bored of them. The only danger is making this ability (to make false images) seem new (absolutely not) or rare (not anymore)!
Re: Imagen, a text-to-image diffusion model
#647Earlier quoted context omitted.
I also worry about the potential to further stifle human creativity, e.g. why paint that oil painting of a panda riding a bicycle when I could generate one in seconds?
One reason: A digital picture of an oil painting != an actual oil painting Of course once someone trains an AI with a robotic arm to do the actual painting, then your worry holds firm.
It's been done, starting from plotter based solutions years ago, through the work of folks like Thomas Lindemeier:
https://scholar.google.com/citations?user=5PpKJ7QAAAAJ&hl=en...
Up to and including actual painting robot arms that dip brushes in paint and apply strokes to canvas today:
https://www.theguardian.com/technology/2022/apr/04/mind-blow...
The painting technique isn't all that great yet for any of these artbots working in a physical medium, but that's largely a general lack of dexterity in manual tool use rather than an art specific challenge. I suspect that RL environments that physically model the application of paint with a brush would help advance the SOTA. It might be cheaper to model other mediums like pencil, charcoal, or even airbrushing first, before tackling more complex and dimensional mediums like oil paint or watercolor.
Re: Imagen, a text-to-image diffusion model
#648Earlier quoted context omitted.
Literally the same thing could be said about Google images, but google images is obviously avaliable to the public. Google knows this will be an unlimited money generator so they're keeping a lid on it.
Given that there's already many competing models in this space prior to any of them having been brought to market, it seems more likely that it will be commoditized.
Re: Imagen, a text-to-image diffusion model
#649Earlier quoted context omitted.
I think they address some of the reasoning behind this pretty clearly in the write-up as well? > The potential risks of misuse raise concerns regarding responsible open-sourcing of code and demos. At this time we have decided not to release code or a public demo. In future work we will explore a framework for responsible externalization that balances the value of external auditing with the risks of unrestricted open-…
I find this is a really bad precedent. Not far from now they'll achieve "super human" level general AI, and be like "yeah it's too powerful for you, we'll keep it internal".
Also, given the processing power and data requirements to create one, there are only a few candidates out there who can get there firstish.