Live data from Hacker News

Imagen, a text-to-image diffusion model

gweb-research-imagen.appspot.com

591–600 of 661 posts

Re: Imagen, a text-to-image diffusion model

#591
post #560

Can anybody give me short high-level explanation how the model achieves these results? I'm especially interested in the image synthesis, not the language parsing. For example, what kind of source images are used for the snake made of corn[0]? It's baffling to me how the corn is mapped to the snake body. [0] https://gweb-research-imagen.appspot.com/main_gallery_images...

> Since guidance weights are used to control image quality and text alignment, we also report ablation results using curves that show the trade-off between CLIP and FID scores as a function of the guidance weights (see Fig. A.5a). We observe that larger variants of T5 encoder results in both better image-text alignment, and image fidelity. This emphasizes the effectiveness of large frozen text encoders for text-to-image models

I usually consider myself fairly intelligent, but I know that when I read an AI research paper I'm going to feel dumb real quick. All I managed to extract from the paper was a) there isn't a clear explanation of how it's done that was written for lay people and b) they are concerned about the quality and biases in the training sets.

Having thought about the problem of "building" an artificial means to visualize from thought, I have a very high level (dumb) view of this. Some human minds are capable of generating synthetic images from certain terms. If I say "visualize a GREEN apple sitting on a picnic table with a checkerboard table cloth", many people will create an image that approximately matches the query. They probably also see a red and white checkerboard cloth because that's what most people have trained their models on in the past. By leaving that part out of the query we can "see" biases "in the wild".

Of course there are people that don't do generative in-mind imagery, but almost all of us do build some type of model in real time from our sensor inputs. That visual model is being continuously updated and is what is perceived by the mind "as being seen". Or, as the Gorillaz put it:

  … For me I say God, y'all can see me now
  'Cos you don't see with your eye
  You perceive with your mind
  That's the end of it…
To generatively produce strongly accurate imagery from text, a system needs enough reference material in the document collection. It needs to have sampled a lot of images of corn and snakes. It needs to be able to do image segmentation and probably perspective estimation. It needs a lot of semantic representations (optimized query of words) of what is being seen in a given image, across multiple "viewing models", even from humans (who also created/curated the collections). It needs to be able to "know" what corn looks like, even from the perspective of another model. It needs to know what "shape" a snake model takes and how combining the bitmask of the corn will affect perspective and framing of the final image. All of this information ends up inside the model's network.

Miika Aittala at Nvidia Research has done several presentations on taking a model (imagined as a wireframe) and then mapping a bitmapped image onto it with a convolutional neural network. They have shown generative abilities for making brick walls that looks real, for example, from images of a bunch of brick walls and running those on various wireframes.

Maybe Imagen is an example of the next step in this, by using diffusion models instead of the CNN for the generator and adding in semantic text mappings while varying the language models weights (i.e. allowing the language model to more broadly use related semantics when processing what is seen in a generated image). I'm probably wrong about half that.

Here's my cut on how I saw this working from a few years ago: https://storage.googleapis.com/mitta-public/generate.PNG

Regardless of how it works, it's AMAZING that we are here now. Very exciting!

Re: Imagen, a text-to-image diffusion model

#592
post #576
post #462

Earlier quoted context omitted.

You can sort of do that with https://fairuseify.ml

I believe that this tech is possible, but this site doesn't provide it. Look at the source of the page: it's just a bunch of sleeps and then you 'download' the same file you provided.

The tech may be possible, but it won't solve anyone's copyright problems. The result would be a "derived work" of the original, irrespective of whether it sounded similar or not.

Re: Imagen, a text-to-image diffusion model

#593

Earlier quoted context omitted.

I think the serious answer is that it is yet another labor multiplier like electricity and software. Our tech since the industrial revolution has allowed us to elevate ourselves from a largely agrarian society to space and cyberspace. AI, by all appearances, continues to be a tool, just the latest in a long line of better tools. It still requires a human to provide intent and direction. Right now in my job, I command…

Well, anyone over 40 will be fucked. There goes your utopia.

No because once this is live, creating private (teaching) assistants and good UX will be cheaper.

Re: Imagen, a text-to-image diffusion model

#594
post #358
post #59

Earlier quoted context omitted.

> the risks of unrestricted open-access What exactly is the risk?

"Make a photograph of Joe Biden in a hotel room bed with Kim Jong-un." Simply the ease at which people are going to be able to make extremely-realistic game photographs is going to do some damage to the world. It's inevitable, but it might be good to postpone it.

The counter argument is that, by the time these models become available to the public, they will produce output that cannot be distinguished from real photos, so the damage will be even greater than if they became available today

Re: Imagen, a text-to-image diffusion model

#596
It’s terrifying that all of these models are one colab notebook away from unleashing unlimited, disastrous imagery on the internet. At least some companies are starting to realize this and are not releasing the source code. However they always manage to write a scientific paper and blog post detailing the exact process to create the model, so it will eventually be recreated by a third party.

Meanwhile, Nvidia sees no problem with yeeting stylegan and and models that allow real humans to be realistically turned into animated puppets in 3d space. The inevitable end result of these scientific achievements will be orders of magnitude worse than deepfakes.

Oh, or a panda wearing sunglasses, in the desert, digital art.

Re: Imagen, a text-to-image diffusion model

#598
post #581

Earlier quoted context omitted.

No

To expand a bit for the grandparent, if you check out this authors other repos you'll notice they have a thing for implementing these papers (multiple DALLE-2 implementations for instance). You should expect to see an implementation there pretty quickly I'd guess.

Not to diminish their contribution but implementing the model is only one third of the battle. The rest is building the training dataset and training the model on a big computer.

Re: Imagen, a text-to-image diffusion model

#599

It’s terrifying that all of these models are one colab notebook away from unleashing unlimited, disastrous imagery on the internet. At least some companies are starting to realize this and are not releasing the source code. However they always manage to write a scientific paper and blog post detailing the exact process to create the model, so it will eventually be recreated by a third party. Meanwhile, Nvidia sees no…

I am absolutely terrified of all this for a different reason: all human professions (not just art) will soon be replaced by “good enough” AI, creating a world flooded with auto-generated junk and billions of people trapped permanently in slums, because you can’t compete with free, and no one can earn a living any longer.

It’s an old fear for sure but it seems to be getting closer and closer every day, and yet most of the discussion around these things seems to be variations of “isn’t this cool?”

Re: Imagen, a text-to-image diffusion model

#600

Earlier quoted context omitted.

Seeing this a lot on youtube also. Scripts pulling in "news" from a source as a script for a robo voice combined with "related" images stitched together randomly.

Even though it's not AI, this is already happening with a lot of content farms. There was a good video a couple years ago from Ann Reason of "How to Cook That" that basically pointed out how the visually-appealing-but-not-actually-feasible "hands and pans" content farms (So Tasty, 5 Minute Crafts, etc.) were killing genuine baking channels. Imagine that instead of having cheap labor from Southeast Asia churn out thes…

> Ann Reason

"Anne Reardon" (autocorrect wah wah waaah)

Post reply on HN