Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

211–220 of 661 posts

Re: Imagen, a text-to-image diffusion model

#211
post #132

Interesting and cool technology - but I can't seem to ignore that every high-quality AI art application is always closed, and I don't seem to buy the ethics excuse for that. The same was said for GPT, yet I see nothing but creativity coming out from its users nowadays.

Granted that's a selection bias: you likely won't hear about the cases where legit obscene output occurs. (the only notable case I've heard is the AI Dungeon incident)

Re: Imagen, a text-to-image diffusion model

#212
post #202

Earlier quoted context omitted.

You mean like: https://say-can.github.io/ This is common in the research PA. People don't want to deal with broccoli man [1]. [1] https://www.youtube.com/watch?v=3t6L-FlfeaI

Looking at that link, I don't think that is a GitHub publication? It is marked Robotics at Google and Everyday Robotics.

My bad, it's a google-specific problem.

Re: Imagen, a text-to-image diffusion model

#213
post #157

Earlier quoted context omitted.

That's an unanswerable question. Perhaps the answer is "don't". Siri takes this approach for a wide range of queries.

How do you pick what should and shouldn't be restricted? Is there some "offense threshold"? I suspect all queries relating to religion, ethnicity, sexuality, and gender will need to be restricted, which almost certainly means you probably can't include humans at all, other than ones artificially inserted with mathematically proven random attributes. Maybe that's why none are in this demo.

"Is Taiwan a country" also comes to mind.

Re: Imagen, a text-to-image diffusion model

#214
post #80

Earlier quoted context omitted.

I think the statistics/representation problem is a big problem on its own, but IMO the bigger problem here is democratizing access to human-like creativity. Currently, the ability to create compelling art is only held by those with some artistic talent. With a tool like this, that restriction is gone. Everyone, no matter how uncreative, untalented, or uncommitted, can create compelling visuals, provided they can use…

> So even if we managed to create a perfect model of representation and inclusion, people could still use it to generate extremely offensive images with little effort. I think people see that as profoundly dangerous. Do they see it as dangerous? Or just offensive? I can understand why people wouldn’t want a tool they have created to be used to generate disturbing, offensive or disgusting imagery. But I don’t really s…

> In fact, I wonder if this sort of technology could reduce the harm caused by people with an interest in disgusting images, because no one needs to be harmed for a realistic image to be created. I am creeping myself out with this line of thinking, but it seems like one potential beneficial - albeit disturbing - outcome.

Interesting idea, but is there any evidence that e.g. consuming disturbing images makes people less likely to act out on disturbing urges? Far from catharsis, I'd imagine consumption of such material to increase one's appetite and likelihood of fulfilling their desires in real life rather than to decrease it.

I suppose it might be hard to measure.

Re: Imagen, a text-to-image diffusion model

#215
post #74

Metacalculus, a mass forecasting site, has steadily brought forward the prediction date for a weakly general AI. Jaw-dropping advances like this, only increase my confidence in this prediction. "The future is now, old man." https://www.metaculus.com/questions/3479/date-weakly-general...

I don't see how this gets us (much) closer to general AI. Where is the reasoning?

Big pretrained models are good enough now that we can pipe them together in really cool ways and our representations of text and images seem to capture what we “mean.”

Re: Imagen, a text-to-image diffusion model

#216
Is there anything at all, besides the training images and labels, that would stop this from generating a convincing response to "A surveillance camera image of Jared Kushner, Vladimir Putin, and Alexandria Ocasio-Cortez naked on a sofa. Jeffrey Epstein is nearby, snorting coke off the back of Elvis"?

Re: Imagen, a text-to-image diffusion model

#217
post #159
post #74

Earlier quoted context omitted.

I don't see how this gets us (much) closer to general AI. Where is the reasoning?

Perhaps the confluence of NLP and something generative?

That doesn’t even lead in the direction of an AGI. The larger and more expensive a model is the less like an “AGI” it is - an independent agent would be able to learn online for free, not need millions in TPU credits to learn what color an apple is.

Re: Imagen, a text-to-image diffusion model

#218
post #188
post #95

Earlier quoted context omitted.

Additionally, if you optimize for most-likely-as-best, you will end up with the stereotypical result 100% of the time, instead of in proportional frequency to the statistics. Put another way, when we ask for an output optimized for "nursiness", is that not a request for some ur stereotypical nurse?

You could stipulate that it roll a die based on percentage results - if 70% of Americans are "white", then 70% of the time show a white person - 13% of the time the result should be black, etc. That's excessively simplified but wouldn't this drop the stereotype and better reflect reality?

No, because a user will see a particular image not the statistically ensemble. It will at times show an Eskimo without a hand because they do statistically exist. But the user definitely does not want that.

Re: Imagen, a text-to-image diffusion model

#219
post #179
post #59

Earlier quoted context omitted.

> the risks of unrestricted open-access What exactly is the risk?

See section 6 titled “Conclusions, Limitations and Societal Impact” in the research paper: https://gweb-research-imagen.appspot.com/paper.pdf One quote: > “On the other hand, generative methods can be leveraged for malicious purposes, including harassment and misinformation spread [20], and raise many concerns regarding social and cultural exclusion and bias [67, 62, 68]”

But do we trust that those who do have access won't be using it for "malicious purposes" (which they might not think is malicious, but perhaps it is to those who don't have access)?

Re: Imagen, a text-to-image diffusion model

#220
post #80

Earlier quoted context omitted.

I think the statistics/representation problem is a big problem on its own, but IMO the bigger problem here is democratizing access to human-like creativity. Currently, the ability to create compelling art is only held by those with some artistic talent. With a tool like this, that restriction is gone. Everyone, no matter how uncreative, untalented, or uncommitted, can create compelling visuals, provided they can use…

> So even if we managed to create a perfect model of representation and inclusion, people could still use it to generate extremely offensive images with little effort. I think people see that as profoundly dangerous. Do they see it as dangerous? Or just offensive? I can understand why people wouldn’t want a tool they have created to be used to generate disturbing, offensive or disgusting imagery. But I don’t really s…

> > ... people could still use it to generate extremely offensive images with little effort. I think people see that as profoundly dangerous. > Do they see it as dangerous? Or just offensive?

I won't speak to whether something is "offensive", but I think that having underlying biases in image-classification or generation has very worrying secondary effects, especially given that organizations like law enforcement want to do things like facial recognition. It's not a perfect analogue, but I could easily see some company pitch a sketch-artist-replacement service that generated images based on someone's description. The potential for having inherent bias present in that makes that kind of thing worrying, especially since the people in charge of buying it are likely to care, or notice, about the caveats.

It does feel like a little bit of a stretch, but at the same time we've also seen such things happen with image classification systems.

Post reply on HN