Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

381–390 of 661 posts

Re: Imagen, a text-to-image diffusion model

#381
post #250
post #132

Interesting and cool technology - but I can't seem to ignore that every high-quality AI art application is always closed, and I don't seem to buy the ethics excuse for that. The same was said for GPT, yet I see nothing but creativity coming out from its users nowadays.

That only lasts until the community copies the paper and catches up. For example the open source DALLE-2 implementation is coming along great: https://github.com/lucidrains/DALLE2-pytorch

Imagen actually shows some of the components in DALLE2 is unnecessary, so Imagen will end up being easier to build. I'll definitely add the dynamic thresholding trick from Imagen to DALLE2 repository though; that is a finding that should boost any DDPMs using classifier free guidance.

Re: Imagen, a text-to-image diffusion model

#382
Probably just a frontend coding mistake, and not an error in the model, but in the interactive example if you select:

"A photo of a Shiba Inu dog Wearing a (sic) sunglasses And black leather jacket Playing guitar In a garden"

The Shiba Inu is not playing a guitar.

Re: Imagen, a text-to-image diffusion model

#384

I thought I was doing well after not being overly surprised by DALL-E 2 or Gato. How am I still not calibrated on this stuff? I know I am meant to be the one who constantly argues that language models already have sophisticated semantic understanding, and that you don't need visual senses to learn grounded world knowledge of this sort, but come on, you don't get to just throw T5 in a multimodal model as-is and have i…

I firmly believe that ~20-40% of the machine learning community will say that all ML models are dumb statistical interpolators all the way until a few years after we achieve AGI. Roughly the same groups will also claim that human intelligence is special magic that cannot be recreated using current technology.

I think it’s in everyone’s benefit if we start planning for a world where a significant portion of the experts are stubbornly wrong about AGI. As a technology, generally intelligent ML has the potential to change so many aspects of our world. The dangers of dismissing the possibility of AGI emerging in the next 5-10 years are huge.

Re: Imagen, a text-to-image diffusion model

#385
post #250

Earlier quoted context omitted.

That only lasts until the community copies the paper and catches up. For example the open source DALLE-2 implementation is coming along great: https://github.com/lucidrains/DALLE2-pytorch

Imagen actually shows some of the components in DALLE2 is unnecessary, so Imagen will end up being easier to build. I'll definitely add the dynamic thresholding trick from Imagen to DALLE2 repository though; that is a finding that should boost any DDPMs using classifier free guidance.

[deleted]

Re: Imagen, a text-to-image diffusion model

#386

Earlier quoted context omitted.

GTP-3 was an erotica virtuoso before it was gagged. There's a serious use case here in endless porn generation. Google would very much like to not be in that business. That said, you can download Dream by Wombo from the app store and it is one of the top smartphone apps, even though it is a few generations behind state of the art.

It is a shame that such powerful AI had to arrive during an age of prudishness that makes the Victorians seem wild.

No post body was provided.

Re: Imagen, a text-to-image diffusion model

#387
post #384

I thought I was doing well after not being overly surprised by DALL-E 2 or Gato. How am I still not calibrated on this stuff? I know I am meant to be the one who constantly argues that language models already have sophisticated semantic understanding, and that you don't need visual senses to learn grounded world knowledge of this sort, but come on, you don't get to just throw T5 in a multimodal model as-is and have i…

I firmly believe that ~20-40% of the machine learning community will say that all ML models are dumb statistical interpolators all the way until a few years after we achieve AGI. Roughly the same groups will also claim that human intelligence is special magic that cannot be recreated using current technology. I think it’s in everyone’s benefit if we start planning for a world where a significant portion of the expert…

> The dangers of dismissing the possibility of AGI emerging in the next 5-10 years are huge.

Again, I think we should consider "The Human Alignment Problem" more in this context. The transformers in question are large, heavy and not really prone to "recursive self-improvement".

If the ML-AGI works out in a few years, who gets to enter the prompts?

Re: Imagen, a text-to-image diffusion model

#388
post #132

Interesting and cool technology - but I can't seem to ignore that every high-quality AI art application is always closed, and I don't seem to buy the ethics excuse for that. The same was said for GPT, yet I see nothing but creativity coming out from its users nowadays.

GTP-3 was an erotica virtuoso before it was gagged. There's a serious use case here in endless porn generation. Google would very much like to not be in that business. That said, you can download Dream by Wombo from the app store and it is one of the top smartphone apps, even though it is a few generations behind state of the art.

"Google would very much like to not be in that business."

Google is not a hobby project anymore: "don't do evil" or whatever they whittered on about back in the day.

Re: Imagen, a text-to-image diffusion model

#389

Earlier quoted context omitted.

I want to be clear, all of this stuff is fascinating, expensive, and difficult. With the possible exception of a few trailer-park weirdos like me, it basically takes a PhD to even stay on top of the field, and you clearly know your stuff. And to be equally clear, I have no inside baseball on how Brain/DM choose when to publish. I have some watercooler chat on the friendly but serious rivalry between those groups, but…

Scaling up improved versions of existing recipes can be done surprisingly fast if you have strong DL infrastructure. Also, GPT-3 was built on top of previous advances such as Google’s BERT. I’m surprised that it took Google so long to answer w/ PaLM, though it seems plausible to me that they wanted a clear enough qualitative advancement that people didn’t immediate say, “So what.” You could’ve had the same reaction y…

There are a lot of really knowledgeable people on here, but this field is near and dear to my heart and it’s obvious that you know it well.

I don’t know what “we should grab a coffee or a beer sometime” means in the hyper-global post-C19 era, but I’d love to speak more on this without dragging a whole HN comment thread through it.

Drop me a line if you’re inclined: ben.reesman at gmail

Re: Imagen, a text-to-image diffusion model

#390
post #132

Interesting and cool technology - but I can't seem to ignore that every high-quality AI art application is always closed, and I don't seem to buy the ethics excuse for that. The same was said for GPT, yet I see nothing but creativity coming out from its users nowadays.

Running inference on one of these models takes like a GPU minute, so they can't just let the public use them.
Post reply on HN