Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

321–330 of 661 posts

Re: Imagen, a text-to-image diffusion model

#322
post #24

Earlier quoted context omitted.

The big labs have become very sensitive with large model releases. It's too easy to make them generate bad PR, to the point of not releasing almost any of them. Flamingo was also a pretty great vison-language model that wasn't released, not even in a demo. PaLM is supposedly better than GPT-3 but closed off. It will probably take a year for open source models to appear.

The largest models which generate the headline benchmarks are never released after any number of years, it seems. Very difficult to replicate results.

[deleted]

Re: Imagen, a text-to-image diffusion model

#323
post #318

Earlier quoted context omitted.

It's an unfortunate reflection of reality. There are three possible outcomes: 1. The model provides a reflection of reality, as politically inconvenient and hurtful as it may be. 2. The model provides an intentionally obfuscated version with either random traits or non correlative traits. 3. The model refuses to answer. Which of these is ideal to you?

What makes you think those are the only options? Why can't we have an option that the model returns a range of different outputs based off a prompt? A model that returns 100% of nurses as female might be statistically more accurate than a model that returns 50% of nurses as female, but it is still not an accurate reflection of the real world. I agree that the model shouldn't return a male nurse 50% of the time. Yet a…

So, the model should have a knowledge of political correctness, and return multiple results if the first choice might reinforce a stereotype?

Re: Imagen, a text-to-image diffusion model

#324

How the fck are things advancing so fast? Is it about to level off …or extend to new domains? What’s a comparable set of technical advances?

Bigger model = better because a lot of performance at this task is memorization or the “lottery ticket hypothesis”.

An impressive advance would be a small model that’s capable of working from an external memory rather than memorizing it.

Re: Imagen, a text-to-image diffusion model

#325

Earlier quoted context omitted.

Figure A.4 in the linked paper is a good high level overview of this model. Shame it was hidden away on page 19 in the appendix! Each box you see there has a section in the paper explaining it in more detail.

Uhh, yeah, I'm going to need much more of an ELI5 than that! Looking at Figure A.4, I understand (again, at a very high-level) the first step of "Frozen Text Encoder", and I have a decent understanding of the upsampling techniques used in the last 2 diffusion model steps, but the middle "Text-to-Image Diffusion Model" step that magically outputs a 64x64 pixel image of an actual golden retriever wearing an actual blue…

A good explanation is here.

https://www.youtube.com/watch?v=344w5h24-h8

Re: Imagen, a text-to-image diffusion model

#326
post #132

Interesting and cool technology - but I can't seem to ignore that every high-quality AI art application is always closed, and I don't seem to buy the ethics excuse for that. The same was said for GPT, yet I see nothing but creativity coming out from its users nowadays.

I don't buy the ethics but I do buy the obvious PR nightmare that would inevitably happen if journalists could play with this and immediately publish their findings of "racist imagery generated by googles AI". That's all it's about and us complaining is not going to make them change their minds.

Re: Imagen, a text-to-image diffusion model

#327
post #314

I apologize in advance for the elitist-sounding tone. In my defense the people I’m calling elite I have nothing to do with, I’m certainly not talking about myself. Without a fairly deep grounding in this stuff it’s hard to appreciate how far ahead Brain and DM are. Neither OpenAI nor FAIR ever has the top score on anything unless Google delays publication . And short of FAIR? D2 lacrosse. There are exceptions to such…

In short, it’s all about money.

Yes and no.

For example: the high-frequency trading industry is estimated to have made somewhere between 2-3 billion dollars in all of 2020, profit/earnings. That’s a good weekend at Google.

HFT shops pay well, but not much different to top performers at FAANG.

People work in HFT because without taking a pay cut they can play real ball: they want to try themselves against the best.

Heavy learning people are no different in wanting both a competitive TC but maybe even more to be where the action is.

That’s currently Blade Runner Industries Ltd, but that could change.

Re: Imagen, a text-to-image diffusion model

#328

Earlier quoted context omitted.

Google is very conservative about anything that can generate open-ended outputs. Also these models are still very expensive computationally.

They're expensive to train, but not awfully expensive to use. Especially if you have hundreds of images you want to generate (due to the way compute devices tend to get much more efficiency with a large batch size). Google could totally afford it, especially if the feature was hidden behind a button the user had to click, and not just run for every image search.

The input control is pretty hard - it kinda needs an AGI :). How do you stop undesirable images being created?

Re: Imagen, a text-to-image diffusion model

#329

I apologize in advance for the elitist-sounding tone. In my defense the people I’m calling elite I have nothing to do with, I’m certainly not talking about myself. Without a fairly deep grounding in this stuff it’s hard to appreciate how far ahead Brain and DM are. Neither OpenAI nor FAIR ever has the top score on anything unless Google delays publication . And short of FAIR? D2 lacrosse. There are exceptions to such…

This characterization is not really accurate. OpenAI has had almost a 2 year lead with GPT-3 dominating the discussion of LLMs (large language models). Google didn’t release its paper on the powerful PaLM-540b model until recently. Similarly, CLiP, Glide, DALL-E, and DALL-E2 have been incredibly influential in visual-language models. Imagen, while highly impressive, definitely is a catch-up piece of work (as was PaLM-540b).

Google clearly demonstrates their unrivaled capability to leverage massive quantities of data and compute, but it’s premature to declare that they’ve secured victory in the AI Wars.

Re: Imagen, a text-to-image diffusion model

#330

Earlier quoted context omitted.

Yes, the idea is that just because it doesn't align to Western ideals of what seems unbiased doesn't mean that the same is necessarily true for other cultures, and by failing to release the model because it doesn't conform to Western, left wing cultural expectations, the authors are ignoring the diversity of cultures that exist globally.

No, it's coming from a perspective of moral realism. It's an objective moral truth that racial and ethnic biases are bad. Yet most cultures around the world are racist to at least some degree, and to they extent that the cultures do, they are bad. The argument you're making, paraphrased, is that the idea that biases are bad is itself situated in particular cultural norms. While that is true to some degree, from a mor…

Western liberal culture says discriminating against one set of minorities to benefit another (affirmative action) is a good thing. What constitutes a racial and ethnic bias is not objective. And therefore Google shouldn't pretend like it is either.

> from a moral realist perspective we can still objectively judge those cultural norms to be better or worse than alternatives

No, because depending on what set of values you have, it is easy to say that one set of biases is better than another. The entire point is that it should not be Google's role to make that judgement - people should be able to do it for themselves.

Post reply on HN