Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

491–500 of 661 posts

Re: Imagen, a text-to-image diffusion model

#491
Seeing the artificial restrictions to this model as well as to DALL-E 2, I can't help but ask myself why the porn industry isn't driving its own research. Given the size of that industry and the sheer abundance of training material, it seems just a matter of time until you can create photo realistic images of yourself with your favourite celebrity for a small fee. Is there anything I am missing? Can you only do this kind of research at google or openai scale?

Re: Imagen, a text-to-image diffusion model

#492
post #384

Earlier quoted context omitted.

I firmly believe that ~20-40% of the machine learning community will say that all ML models are dumb statistical interpolators all the way until a few years after we achieve AGI. Roughly the same groups will also claim that human intelligence is special magic that cannot be recreated using current technology. I think it’s in everyone’s benefit if we start planning for a world where a significant portion of the expert…

You should be much more concerned about the prospect of nuclear war right now than the sudden emergence of an AGI.

100 times this. There’s very little sign of AGI, but nuclear weapons exist, can definitely destroy the planet already, are designed to, have nearly done so in the past, and we’re at the most dangerous point in decades.

Re: Imagen, a text-to-image diffusion model

#493
post #437

I have to wonder how much releasing these models will "poison the well" and fill the internet with AI generated images that make training an improved model difficult. After all if every 9/10 "oil painted" image online starts being from these generative models it'll become increasingly difficult to scrape the web and to learn from real world data in a variety of domains. Essentially once these things are widely availa…

People training newer models just have to look for the "Imagen" tag or the Dall-E2 rainbow at the corner and heuristically exclude images having these. This is trivial.

Unless you assume there are bad actors who will crop out the tags. Not many people now have access to Dall-E2 or will have access to Imagen.

As someone working in Vision, I am also thinking about whether to include such images deliberately. Using image augmentation techniques is ubiquitous in the field. Thus we introduce many examples for training the model that are not in the distribution over input images. They improve model generality by huge margins. Whether generated images improve generality of future models is a thing to try.

Damn I just got an idea for a paper writing this comment.

Re: Imagen, a text-to-image diffusion model

#494
post #491

Seeing the artificial restrictions to this model as well as to DALL-E 2, I can't help but ask myself why the porn industry isn't driving its own research. Given the size of that industry and the sheer abundance of training material, it seems just a matter of time until you can create photo realistic images of yourself with your favourite celebrity for a small fee. Is there anything I am missing? Can you only do this…

Porn is actually a really good litmus test to see if a money/media transfer technology has real promise. Pornography needs exactly 2 things to work well - a way to deliver media, and a way to collect money. If you truly have a system that can do one of those two things better than we currently can, and it's not just empty hype, it will be used for porn. "Empty hype" won't touch that stuff, but real-world usecases will.

Unrelated to the main topic, but this is exactly why I think cryptocurrencies will only be used for illegal activities, or things you may want to hide, and nothing else. Because that's where it has found its usecase in porn.

Re: Imagen, a text-to-image diffusion model

#496
post #491

Seeing the artificial restrictions to this model as well as to DALL-E 2, I can't help but ask myself why the porn industry isn't driving its own research. Given the size of that industry and the sheer abundance of training material, it seems just a matter of time until you can create photo realistic images of yourself with your favourite celebrity for a small fee. Is there anything I am missing? Can you only do this…

Transfer learning is a thing.

But I have not tried making generative models with out-of-distribution data before. Distributions other than main training data.

There are several indie attempts that I am aware of. Mentioning them to the reply of this comment. (In case the comment gets deleted)

The first layers should be general. But the later layers should not behave well to porn images. As they are more specialist layers learning distribution specific visual patters.

Transfer learning is posssible.

Re: Imagen, a text-to-image diffusion model

#497
post #473

One thing that no one predicted in AI development was how good it would become at some completely unexpected tasks while being not so great at the ones we supposed/hoped it would be good. AI was expected to grow like a child. Somehow blurting out things that would show some increasing understanding on a deep level but poor syntax. In fact we get the exact opposite. AI is creating texts that are syntaxically correct a…

I doubt 99% of humans can draw a ”chess game with a puzzle where white mates in 4 moves”

Re: Imagen, a text-to-image diffusion model

#498

Earlier quoted context omitted.

GTP-3 was an erotica virtuoso before it was gagged. There's a serious use case here in endless porn generation. Google would very much like to not be in that business. That said, you can download Dream by Wombo from the app store and it is one of the top smartphone apps, even though it is a few generations behind state of the art.

"Google would very much like to not be in that business." Google is not a hobby project anymore: "don't do evil" or whatever they whittered on about back in the day.

I imagine it's viewing porn not as "evil" but as "something we should absolutely never even come close to touching or talking about and is best left as something we pretend doesn't exist so we don't get regulated out of existence"

Re: Imagen, a text-to-image diffusion model

#499

Earlier quoted context omitted.

"Reality" as defined by the available training set isn't necessarily reality. For example, Google's image search results pre-tweaking had some interesting thoughts on what constitutes a professional hairstyle, and that searches for "men" and "women" should only return light-skinned people: https://www.theguardian.com/technology/2016/apr/08/does-goog... Does that reflect reality? No. (I suspect there are also mostly u…

You know, it wouldn't surprise me if people talking about how black curly hair shouldn't be seen as unprofessional contributed to google thinking there's an association between the concepts of "unprofessional hair" and "black curly hair"

That's exactly what's happening. Doing the search from the article of "unprofessional hair for work" brings up images with headlines like "It's ridiculous to say that black women's hair is unprofessional". (In addition to now bringing up images from that article itself and other similar articles comparing Google Images searches.)

Re: Imagen, a text-to-image diffusion model

#500
post #473

One thing that no one predicted in AI development was how good it would become at some completely unexpected tasks while being not so great at the ones we supposed/hoped it would be good. AI was expected to grow like a child. Somehow blurting out things that would show some increasing understanding on a deep level but poor syntax. In fact we get the exact opposite. AI is creating texts that are syntaxically correct a…

I doubt 99% of humans can draw a ”chess game with a puzzle where white mates in 4 moves”

Maybe not draw, but we can do an image search for "chess puzzle mate in 4" which gives plenty of results:

https://www.google.com/search?q=chess+puzzle+mate+in+4&tbm=i...

It would be surprising if AI couldn't do the same search and produce a realistic drawing out of any one of the result puzzles.

Post reply on HN