Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

411–420 of 661 posts

Re: Imagen, a text-to-image diffusion model

#411

Probably just a frontend coding mistake, and not an error in the model, but in the interactive example if you select: "A photo of a Shiba Inu dog Wearing a (sic) sunglasses And black leather jacket Playing guitar In a garden" The Shiba Inu is not playing a guitar.

Found the QA tester.

Also, no sunglasses in "A photo of a raccoon wearing sunglasses and a red shirt riding a bike in a garden," and a few similar prompts (e.g. surfing).

Re: Imagen, a text-to-image diffusion model

#412
post #373

Earlier quoted context omitted.

> these networks are really really big, so they can memorize stuff good. People tend to really underestimate just how big these models are. Of course these models aren't simply "really really big" MLPs, but the cleverness of the techniques used to build them is only useful at insanely large scale. I do find these models impressive as examples of "here's what the limit of insane amounts of data, insane amounts of comp…

The hackers will not be far behind. You can run some of the v1 diffusion models on a local machine. I think it's fair to say that this is the way it's always been. In 1990, you couldn't hack on an accurate fluid simulation at home, you needed to be at a university or research lab with access to a big cluster. But then, 10 years later, you could do it on a home PC. And then, 10 years after that, you could do it in a b…

And the supply crunch is getting better (you can buy an RTZ 3080 at MSRP now!) and technological progress doesn't seem to be slowing down. If the rumors are to be believed, a 4090 will be close to twice as fast as a 3090.

Re: Imagen, a text-to-image diffusion model

#413

Earlier quoted context omitted.

> it will be used in ways we haven’t anticipated Oh yeah, as a woman who grew up in a Third World country, how an AI model generates images would have deeply affected my daily struggles! /s It's kinda insulting that they think that this would be insulting. Like "Oh no I asked the model to draw a doctor and it drew a male doctor, I guess there's no point in me pursuing medical studies" ...

Yes actually, subconscious bias due to historical prejudice does have a large effect on society. Obviously there are things with much larger effects, that doesn't mean that this doesn't exist. > Oh no I asked the model to draw a doctor and it drew a male doctor, I guess there's no point in me pursuing medical studies If you don't think this is a real thing that happens to children you're not thinking especially hard.…

> Yes actually, subconscious bias due to historical prejudice does have a large effect on society.

The evidence for implicit bias is pretty weak and IIRC is better explained by people having explicit bias but lying about it when asked.

(Note: this is even worse.)

Re: Imagen, a text-to-image diffusion model

#414

Earlier quoted context omitted.

"Reality" as defined by the available training set isn't necessarily reality. For example, Google's image search results pre-tweaking had some interesting thoughts on what constitutes a professional hairstyle, and that searches for "men" and "women" should only return light-skinned people: https://www.theguardian.com/technology/2016/apr/08/does-goog... Does that reflect reality? No. (I suspect there are also mostly u…

You know, it wouldn't surprise me if people talking about how black curly hair shouldn't be seen as unprofessional contributed to google thinking there's an association between the concepts of "unprofessional hair" and "black curly hair"

You really are not helping that cause.

As a foreigner[], your point confused me anyway, and doing a Google for cultural stuff usually gets variable results. But I did laugh at many of the comments here https://www.reddit.com/r/TooAfraidToAsk/comments/ufy2k4/why_...

[] probably, New Zealand, although foreigner is relative

Re: Imagen, a text-to-image diffusion model

#415

Earlier quoted context omitted.

> But the incentives are weird here: who, exactly, does it serve for us plebs to be able to train these things from scratch? I'm not sure it matters. The history of computing shows that within the decade we will all have the ability to train and use these models.

... Unless the possession of capable models becomes a legal liability by that time.

This won’t happen in an interesting way. What will happen is you’ll find out training a model on copyrighted inputs causes it to memorize those inputs and the owners own your output.

Re: Imagen, a text-to-image diffusion model

#416

I apologize in advance for the elitist-sounding tone. In my defense the people I’m calling elite I have nothing to do with, I’m certainly not talking about myself. Without a fairly deep grounding in this stuff it’s hard to appreciate how far ahead Brain and DM are. Neither OpenAI nor FAIR ever has the top score on anything unless Google delays publication . And short of FAIR? D2 lacrosse. There are exceptions to such…

> Neither OpenAI nor FAIR ever has the top score on anything unless Google delays publication. This is ... very incorrect. I am very certain (95%+) that Google had nothing even close to GPT-3 at the time of its release. It's been 2 full years since GPT-3 was released, and even longer since OpenAI actually trained it. That's not to talk about any of the other things OpenAI/FAIR has released that were SOTA at the time…

Any “brash generalization” is clearly going to be grossly incorrect in concrete cases, and while I have a little gossip from true insiders, it’s nowhere near enough to make definitive statements about specific progress on teams at companies that I’ve never worked for.

I did a bit of disclaimer on my original post but not enough to withstand detailed scrutiny. This is sort of the trouble with trying to talk about cutting-edge research in what amounts to a tweet: what’s the right amount of oversimplified, emphatic statement to add legitimate insight but not overstep into being just full of shit.

I obviously don’t know that publication schedules at heavy-duty learning shops are deliberate and factor-in other publications. The only one I know anything concretely about is FAIR and even that’s badly dated knowledge.

I was trying to squeeze into a few hundred characters my very strong belief that Brain and DM haven’t let themselves be scooped since ResNet, based on my even stronger belief that no one has the muscle to do it.

To the extent that my oversimplification detracted from the conversation I regret that.

Re: Imagen, a text-to-image diffusion model

#418
post #224

Earlier quoted context omitted.

No, it's coming from a perspective of moral realism. It's an objective moral truth that racial and ethnic biases are bad. Yet most cultures around the world are racist to at least some degree, and to they extent that the cultures do, they are bad. The argument you're making, paraphrased, is that the idea that biases are bad is itself situated in particular cultural norms. While that is true to some degree, from a mor…

You're confused by the double meaning of the word "bias". Here we mean mathematical biases. For example, a good mathematical model will correctly tell you that people in Japan (geographical term) are more likely to be Japanese (ethnic / racial bias). That's not "objectively morally bad", but instead, it's "correct".

Well that's not the issue here, the problem is the examples like searches for images of "unprofessional hair" returning mostly Black people in the results. That is something we can judge as objectively morally bad.

Re: Imagen, a text-to-image diffusion model

#419

Is there anything at all, besides the training images and labels, that would stop this from generating a convincing response to "A surveillance camera image of Jared Kushner, Vladimir Putin, and Alexandria Ocasio-Cortez naked on a sofa. Jeffrey Epstein is nearby, snorting coke off the back of Elvis"?

- The current examples aren’t convincing pictures of “a shiba inu playing a guitar”.

- If you made that picture with actors or in MS Paint, politics boomers on Facebook wouldn’t care either way. They’d just start claiming it’s real if they like the message.

Re: Imagen, a text-to-image diffusion model

#420

Earlier quoted context omitted.

It’s wild to me that the HN consensus is so often that 1) discourse around the internet is terrible, it’s full of spam and crap, and the internet is an awful unrepresentative snapshot of human existence, and 2) the biases of general-internet-training-data are fine in ML models because it just reflects real life.

The bias on HN is that people who prioritize being nice, or may possibly have humanities degrees or be ultra-libs from SF, are wrong because the correct answer would be cynical and cold-heartedly mechanical. Other STEM adjacent communities feel similarly but I don’t get it from actual in person engineers much.

Being nice is alright, but why is it that this fundamental drive is so often an uninspiring explanation behind yet another incursion towards one's individual freedom, even if exercising said freedom doesn't bring any real harm to anyone involved?

Maybe the engineers conclude correctly that voicing this concern without the veil of anonymity will do nothing good to their humble livelihood, and thus you don't hear it from them in person.

Post reply on HN