Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

271–280 of 661 posts

Re: Imagen, a text-to-image diffusion model

#271

I give it a few years before Google makes stock images irrelevant.

Tbh imagine this tech combines particularly well with really well curated stock image databases so outputs can be made with recognisable styles, and actors and design elements can be reused across multiple generated images.

If Getty et al aren't already spending money on that possibility, they probably should be.

Re: Imagen, a text-to-image diffusion model

#272
post #223

Earlier quoted context omitted.

Big pretrained models are good enough now that we can pipe them together in really cool ways and our representations of text and images seem to capture what we “mean.”

Yeah, it seems like it. But it's still just complicated statistical models. Again, where is the reasoning?

I still think we're missing some fundamental insights on how layered planning/forecasting/deducting/reasoning works, and that figuring this out will be necessary in order to create AI that we could say "reasons".

But with the recent advances/demonstrations, it seems more likely today than in 2019 that our current computational resources are sufficient to perform magnificantly spooky stuff if they're used correctly. They are doing that already already, and that's without deliberately making the software do anything except draw from a vast pool of examples.

I think it's reasonable, based on this, to update one's expectations of what we'd be able to do if we figured out ways of doing things that aren't based on first seeing a hundred million examples of what we want the computer to do.

Things that do this can obviously exist, we are living examples. Does figuring it out seem likely to be many decades away?

Re: Imagen, a text-to-image diffusion model

#273
I apologize in advance for the elitist-sounding tone. In my defense the people I’m calling elite I have nothing to do with, I’m certainly not talking about myself.

Without a fairly deep grounding in this stuff it’s hard to appreciate how far ahead Brain and DM are.

Neither OpenAI nor FAIR ever has the top score on anything unless Google delays publication. And short of FAIR? D2 lacrosse. There are exceptions to such a brash generalization, NVIDIA’s group comes to mind, but it’s a very good rule of thumb. Or your whole face the next time you are tempted to doze behind the wheel of a Tesla.

There are two big reasons for this:

- the talent wants to work with the other talent, and through a combination of foresight and deep pockets Google got that exponent on their side right around the time NVIDIA cards started breaking ImageNet. Winning the Hinton bidding war clinched it.

- the current approach of “how many Falcon Heavy launches worth of TPU can I throw at the same basic masked attention with residual feedback and a cute Fourier coloring” inherently favors deep pockets, and obviously MSFT, sorry OpenAI has that, but deep pockets also non-linearly scale outcomes when you’ve got in-house hardware for multiply-mixed precision.

Now clearly we’re nowhere close to Maxwell’s Demon on this stuff, and sooner or later some bright spark is going to break the logjam of needing 10-100MM in compute to squeeze a few points out of a language benchmark. But the incentives are weird here: who, exactly, does it serve for us plebs to be able to train these things from scratch?

Re: Imagen, a text-to-image diffusion model

#275
post #223

Earlier quoted context omitted.

Big pretrained models are good enough now that we can pipe them together in really cool ways and our representations of text and images seem to capture what we “mean.”

Yeah, it seems like it. But it's still just complicated statistical models. Again, where is the reasoning?

All it takes is one 'trick' to give these models the ability to do reasoning.

Like for example the discovery that language models get far better at answering complex questions if asked to show their working step by step with chain of thought reasoning as in page 19 of the PaLM paper [1]. Worth checking out the explanations of novel jokes on page 38 of the same paper. While it is, like you say, all statistics, if it's indistinguishable from valid reasoning, then perhaps it doesn't matter.

[1]: https://arxiv.org/pdf/2204.02311.pdf

Re: Imagen, a text-to-image diffusion model

#277
post #269

Earlier quoted context omitted.

what if I asked the model to show me a sunday school photograph of baptists in the National Baptist Convention?

The pictures I got from a similar model when asking for a "sunday school photograph of baptists in the National Baptist Convention": https://ibb.co/sHGZwh7

and how do we _feel_ about that outcome?

Re: Imagen, a text-to-image diffusion model

#278

As someone who has a layman's understanding of neural networks, and who did some neural network programming ~20 years ago before the real explosion of the field, can someone point to some resources where I can get a better understanding about how this magic works? I mean, from my perspective, the skill in these (and DALL-E's) image reproductions is truly astonishing. Just looking for more information about how the so…

> I mean, from my perspective, the skill in these (and DALL-E's) image reproductions is truly astonishing.

A basic part of it is that neural networks combine learning and memorizing fluidly inside them, and these networks are really really big, so they can memorize stuff good.

So when you see it reproduce a Shiba Inu well, don’t think of it as “the model understands Shiba Inus”. Think of it as making a collage out of some Shiba Inu clip art it found on the internet. You’d do the same if someone asked you to make this image.

It’s certainly impressive that the lighting and blending are as good as they are though.

Re: Imagen, a text-to-image diffusion model

#279
post #150

Earlier quoted context omitted.

Not really; the gender of a nurse is accidental, other properties are essential.

How do you know this? Because you can, in your mind, divide the function of a nurse from the statistical reality of nursing? Are the logical divisions you make in your mind really indicative of anything other than your arbitrary personal preferences?

No, because there's at least one male nurse.

Re: Imagen, a text-to-image diffusion model

#280
post #132

Interesting and cool technology - but I can't seem to ignore that every high-quality AI art application is always closed, and I don't seem to buy the ethics excuse for that. The same was said for GPT, yet I see nothing but creativity coming out from its users nowadays.

GTP-3 was an erotica virtuoso before it was gagged. There's a serious use case here in endless porn generation. Google would very much like to not be in that business. That said, you can download Dream by Wombo from the app store and it is one of the top smartphone apps, even though it is a few generations behind state of the art.

The current GPT3 on the OpenAI dashboard is perfectly capable of generating erotica or being racist even if you don’t want it to. They didn’t block it so much as put up a warning dialog when it’s acting up.

Actually, I think they made InstructGPT even better at erotica because it’s trained to be “helpful and friendly”, so in other words they made it a sub.

Post reply on HN