Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

361–370 of 661 posts

Re: Imagen, a text-to-image diffusion model

#361

I apologize in advance for the elitist-sounding tone. In my defense the people I’m calling elite I have nothing to do with, I’m certainly not talking about myself. Without a fairly deep grounding in this stuff it’s hard to appreciate how far ahead Brain and DM are. Neither OpenAI nor FAIR ever has the top score on anything unless Google delays publication . And short of FAIR? D2 lacrosse. There are exceptions to such…

Is Maxwell's Demon applicable to this scenario? I'm not a physicist but I recently had to look it up after talking with someone and thought it had to do with a specific thermodynamic thought experiment with gas particles and heat differences. Is there is another application I don't understand with computing power?

Re: Imagen, a text-to-image diffusion model

#362
post #132

Interesting and cool technology - but I can't seem to ignore that every high-quality AI art application is always closed, and I don't seem to buy the ethics excuse for that. The same was said for GPT, yet I see nothing but creativity coming out from its users nowadays.

GTP-3 was an erotica virtuoso before it was gagged. There's a serious use case here in endless porn generation. Google would very much like to not be in that business. That said, you can download Dream by Wombo from the app store and it is one of the top smartphone apps, even though it is a few generations behind state of the art.

It is a shame that such powerful AI had to arrive during an age of prudishness that makes the Victorians seem wild.

Re: Imagen, a text-to-image diffusion model

#363

I apologize in advance for the elitist-sounding tone. In my defense the people I’m calling elite I have nothing to do with, I’m certainly not talking about myself. Without a fairly deep grounding in this stuff it’s hard to appreciate how far ahead Brain and DM are. Neither OpenAI nor FAIR ever has the top score on anything unless Google delays publication . And short of FAIR? D2 lacrosse. There are exceptions to such…

> But the incentives are weird here: who, exactly, does it serve for us plebs to be able to train these things from scratch? I'm not sure it matters. The history of computing shows that within the decade we will all have the ability to train and use these models.

... Unless the possession of capable models becomes a legal liability by that time.

Re: Imagen, a text-to-image diffusion model

#364

I apologize in advance for the elitist-sounding tone. In my defense the people I’m calling elite I have nothing to do with, I’m certainly not talking about myself. Without a fairly deep grounding in this stuff it’s hard to appreciate how far ahead Brain and DM are. Neither OpenAI nor FAIR ever has the top score on anything unless Google delays publication . And short of FAIR? D2 lacrosse. There are exceptions to such…

Is Maxwell's Demon applicable to this scenario? I'm not a physicist but I recently had to look it up after talking with someone and thought it had to do with a specific thermodynamic thought experiment with gas particles and heat differences. Is there is another application I don't understand with computing power?

You’re absolutely right that I used a sloppy analogy there.

It’s Boltzmann and Szilard that did the original “kT” stuff around underlying thermodynamics governing energy dissipation in these scenarios, and Rolf Landaeur (I think that’s how you spell it) who did the really interesting work on how to apply that thermo work to lower-bounds on energy-expenditure in a given computation.

I said Maxwell’s Demon because it’s the best known example of a deep connection between useful work and computation. But it was sloppy.

Re: Imagen, a text-to-image diffusion model

#365

I apologize in advance for the elitist-sounding tone. In my defense the people I’m calling elite I have nothing to do with, I’m certainly not talking about myself. Without a fairly deep grounding in this stuff it’s hard to appreciate how far ahead Brain and DM are. Neither OpenAI nor FAIR ever has the top score on anything unless Google delays publication . And short of FAIR? D2 lacrosse. There are exceptions to such…

> Neither OpenAI nor FAIR ever has the top score on anything unless Google delays publication.

This is ... very incorrect. I am very certain (95%+) that Google had nothing even close to GPT-3 at the time of its release. It's been 2 full years since GPT-3 was released, and even longer since OpenAI actually trained it.

That's not to talk about any of the other things OpenAI/FAIR has released that were SOTA at the time of release (Dall-E 1, JukeBox, Poker, Diplomacy, Codex).

Google Brain and Deepmind have done a lot of great work, but to imply that they essentially have a monopoly on SOTA results and all SOTA results other labs have achieved are just due to Google delaying publication is ridiculous.

Re: Imagen, a text-to-image diffusion model

#366
post #7

>While we leave an in-depth empirical analysis of social and cultural biases to future work, our small scale internal assessments reveal several limitations that guide our decision not to release our model at this time. Some of the reasoning: >Preliminary assessment also suggests Imagen encodes several social biases and stereotypes, including an overall bias towards generating images of people with lighter skin tones…

> Really sad that breakthrough technologies are going to be withheld due to our inability to cope with the results.

Indeed it is. Consider this an early, toy version of the political struggle related to ownership of AI-scientists and AI-engineers of the near future. That is, generally capable models.

I do think the public should have access to this technology, given so much is at stake. Or at least the scientists should be completely, 24/7, open about their R&D. Every prompt that goes into these models should be visible to everyone.

Re: Imagen, a text-to-image diffusion model

#367
I thought I was doing well after not being overly surprised by DALL-E 2 or Gato. How am I still not calibrated on this stuff? I know I am meant to be the one who constantly argues that language models already have sophisticated semantic understanding, and that you don't need visual senses to learn grounded world knowledge of this sort, but come on, you don't get to just throw T5 in a multimodal model as-is and have it work better than multimodal transformers! VLM[1] at least added fine-tuned internal components.

Good lord we are screwed. And yet somehow I bet even this isn't going to kill off the they're just statistical interpolators meme.

[1] https://www.deepmind.com/blog/tackling-multiple-tasks-with-a...

Re: Imagen, a text-to-image diffusion model

#368

Earlier quoted context omitted.

That’s pretty speculative and dubious (the holding back part) given the heavy bias to publication culture at Google Research and DeepMind. OpenAI has hardly been “crushed” here; PaLM and Imagen are solid, incremental advances, but given what came before them, not Earth-shattering. If I were going to cite evidence for Alphabet’s “supremacy” in AI, I would’ve picked something more novel and surprising such as AlphaFold…

I want to be clear, all of this stuff is fascinating, expensive, and difficult. With the possible exception of a few trailer-park weirdos like me, it basically takes a PhD to even stay on top of the field, and you clearly know your stuff. And to be equally clear, I have no inside baseball on how Brain/DM choose when to publish. I have some watercooler chat on the friendly but serious rivalry between those groups, but…

Scaling up improved versions of existing recipes can be done surprisingly fast if you have strong DL infrastructure. Also, GPT-3 was built on top of previous advances such as Google’s BERT. I’m surprised that it took Google so long to answer w/ PaLM, though it seems plausible to me that they wanted a clear enough qualitative advancement that people didn’t immediate say, “So what.”

You could’ve had the same reaction years ago when Google published GoogleNet followed by a series of increasingly powerful Inception models - namely that Google would wind up owning the DNN space. But it didn’t play out that way, perhaps because Google dragged its feet releasing the models and training code, and by the time it did, there were simpler and more powerful models available like ResNet.

Meta’s recent release of the actual OPT LLM weights is probably going to have more impact than PaLM, unless Google can be persuaded to open up that model.

Re: Imagen, a text-to-image diffusion model

#369

Earlier quoted context omitted.

Is Maxwell's Demon applicable to this scenario? I'm not a physicist but I recently had to look it up after talking with someone and thought it had to do with a specific thermodynamic thought experiment with gas particles and heat differences. Is there is another application I don't understand with computing power?

You’re absolutely right that I used a sloppy analogy there. It’s Boltzmann and Szilard that did the original “kT” stuff around underlying thermodynamics governing energy dissipation in these scenarios, and Rolf Landaeur (I think that’s how you spell it) who did the really interesting work on how to apply that thermo work to lower-bounds on energy-expenditure in a given computation. I said Maxwell’s Demon because it’s…

OK thanks I figured there was a connection between computational power and thermodynamics when you get to a small enough scale but I wasn't sure how to apply it!

Re: Imagen, a text-to-image diffusion model

#370

Interesting to me that this one can draw legible text. DALLE models seem to generate weird glyphs that only look like text. The examples they show here have perfectly legible characters and correct spelling. The difference between this and DALLE makes me suspicious / curious. I wish I could play with this model.

Imagen takes text embeddings, OpenAI model takes image embeddings instead, this is the reason. There are other models that can generate text: latent diffusion trained on LAION-400M, GLIDE, DALL-E (1).
Post reply on HN