Live data from Hacker News

Imagen: An AI system that creates photorealistic images from input text

imagen.research.google

61–70 of 233 posts

Re: Imagen: An AI system that creates photorealistic images from input text

#61

There are some obvious mistakes in tools like this. Such as: Human faces are wrong, writing is usually scrambled, fingers look weird etc... Do you know if we need to have a major breakthrough similar to what happened 6 months ago to fix this or could these be fixed with incremental improvements in current techniques / datasets?

> Human faces are wrong, writing is usually scrambled, fingers look weird

I think a lot of this has been solved in DALL-E already. It's pretty good at right-looking faces, and fingers. Text not so much... but it does appear that whatever OpenAI are doing behind the scenes, that's getting better too.

Re: Imagen: An AI system that creates photorealistic images from input text

#62
post #35

Earlier quoted context omitted.

Like anyone deeply in a field I know maybe several thousand people who could probably give a better answer, but I figure I'll give an effort to provide one since I don't see any good ones posted yet. The moment everyone knew this was going to be big was in 2019 when StyleGAN came out. They used a lot of tricks like aligning face features (like eyes) and had all their pictures of a single domain (the most famous being…

Thanks for answering. Since you mentioned your work on text-to-3d, what are the ways to enhance the image/3d model to actually be photo-(or rather reality)-realistic? Even (presumably) hand-picked examples from google on the linked page lack support bars of the sunglasses, include floating cups of wine with base-less Eiffel tower in the background. P.S. It seems raccoons are unimaginable (even for AI) with any sungla…

I know as much about how to get the best image outputs from text inputs as the person who designed an airport knows the best place to eat in it. The emergent properties of the system are a result of the data put into it, so I can only discuss the system itself, not what it ended up doing with the data in that system.

The models are a product of their datasets, specifically the relationship of the images and prompts via CLIP. CLIP puts both images and text into coordinate space, imagine just a 2D graph. It tries to assure that for any real image and its caption, they will each be each others closest neighbor in that coordinate space.

So if you want a certain image, you have to ask "what caption would be most likely and most uniquely given to the image I'm imagining".

I'm sure this advice is way less helpful than what you find in prompt engineering discord channels and guides I've seen.

Re: Imagen: An AI system that creates photorealistic images from input text

#64

i only have a passing curiosity in these projects personally. can someone in the field explain why this has exploded recently? there seems to be a lot of these tools released recently (text to image) was there a major breakthrough? a new idea that pushed everyone forward? a recent sharing of talent between groups? edit: just another thought, are they just being posted to HN now, i don't see a date on the page for whe…

Like anyone deeply in a field I know maybe several thousand people who could probably give a better answer, but I figure I'll give an effort to provide one since I don't see any good ones posted yet. The moment everyone knew this was going to be big was in 2019 when StyleGAN came out. They used a lot of tricks like aligning face features (like eyes) and had all their pictures of a single domain (the most famous being…

Interesting!

I knew about transformers, CLIP and diffusion, but pixel patch encodings are new to me.

Can you give me more details / point me towards an explainer? A quick duckduckgo search didn't help.

Re: Imagen: An AI system that creates photorealistic images from input text

#65

Earlier quoted context omitted.

Lot of potential for movies- using AI to up res, AI to turn a 2d movie 3d, more advanced would be new or edited scenes. Or how about translating a live action movie to a cartoon and vice versa. Or a different style or tone. Run The Lord of the Rings through a cyberpunk filter.

I'm eager for someone to start recreating the missing episodes of Doctor Who. The audio still exists and there are publicity shots from many of the missing episodes

I wonder how good AI will be at recreating the look of the props made out of broom sticks, gaffer tape and paper mache.

Re: Imagen: An AI system that creates photorealistic images from input text

#66

Smart for Dalle, Midjourney, and Stable Diffusion to capitalize on this quickly. It looks like the technology is being commoditized at rapid speed. I wonder what’s next.

The AI is moving up the "content ladder" generating text -> now images -> next videos

Smart for Google to invest in this because their business relies on third party content to exist (blogs, webpages, youtube videos)

If they can vertically integrate their business to create the content AND own the discovery algorithms then they officially win the internet

Re: Imagen: An AI system that creates photorealistic images from input text

#68

GPT-3(4), DALL-E, Imagen, ... I wonder what's next?

Sound and moving images. Put it all together and fan fiction will never be the same again!

But seriously it may usher in a new cottage industry of content creation, from graphic novels,to animation to movies.

Re: Imagen: An AI system that creates photorealistic images from input text

#69
post #15
post #8

From the cherry-picked example-images on that page, it seems like Imagen more closely follows the prompt than the open Stable Diffusion model[0]. Stable Diffusion needs a lot of hints before it makes out of the ordinary pictures. In general, I think these models are a great and funny toy, but not a threat to stock-photos yet. This may change within a year or three years though. [0]: https://stability.ai/blog/stable-d…

Not a threat to stock photos? That's exactly what they are. Look at these photos from the Midjourney Discord today: crystal dragon thing: https://cdn.discordapp.com/attachments/951197655021797436/10... https://cdn.discordapp.com/attachments/951197655021797436/10... https://cdn.discordapp.com/attachments/951197655021797436/10... davinci-style notebook of flying machines: https://cdn.discordapp.com/attachments/10080491…

They all look great! But the usual customers of stock photos are not looking for dragons or cavemen taking a group selfie. And all the images are not the result of a beginner trying their first prompt, they all took many tries to be generated.

Re: Imagen: An AI system that creates photorealistic images from input text

#70
post #63

I feel like Imagen gets all the noise and people forget about Parti - https://parti.research.google/ It's another google project using a different set of techniques.

I just ignore Google projects since they don’t get released… and when they are released it’s with such severe limitations that they don’t work well.
Post reply on HN