Live data from Hacker News

Google Imagen 2

cloud.google.com

11–20 of 194 posts

Re: Google Imagen 2

#11
post #9
post #3

This would have been an epic release two years ago, but there are now many well-established models in this area (DALL-E, Midjourney, Stable Diffusion). It would be great to see some comparisons or benchmarks to show Imagen 2 is a better alternative. As it stands, it's hard for me to tell if this is worth switching to.

> it's hard for me to tell I can only compare it to Stable Diffusion. But Imagen2 seems significant more advanced. Try to do anything with text and SDxl. It's not easy and often messes up. I don't think you can get a clean logo with multiple text areas on sdxl. Look at the prompt and image of the robin. That is mighty impressive.

Stability AI has gaps in SDXL for text, but they seem to do a better job with Deep Floyd ( https://github.com/deep-floyd/IF ). I have done a lot of interesting text things with Deep Floyd

Re: Google Imagen 2

#12
post #9
post #3

This would have been an epic release two years ago, but there are now many well-established models in this area (DALL-E, Midjourney, Stable Diffusion). It would be great to see some comparisons or benchmarks to show Imagen 2 is a better alternative. As it stands, it's hard for me to tell if this is worth switching to.

> it's hard for me to tell I can only compare it to Stable Diffusion. But Imagen2 seems significant more advanced. Try to do anything with text and SDxl. It's not easy and often messes up. I don't think you can get a clean logo with multiple text areas on sdxl. Look at the prompt and image of the robin. That is mighty impressive.

yeah stable diffusion has very limited understanding of composition instructions. you can reliably get things drawn, but it's super hard to get a specific thing in a specific place (i.e "a man with blonde hairs near a girl with black hairs" is gonna assign hair color more or less randomly and there's no guarantee on how many people will be on the picture) - regional prompting and control net somewhat help, but regional prompting is very unreliable and control net is, well, not text to image.

dalle 3 gets things right most of the time

Re: Google Imagen 2

#13
post #7
post #2

This post has more information: https://cloud.google.com/blog/products/ai-machine-learning/i... I can't figure out how to try this thing. The closest I got was this sentence: "To get started with Imagen 2 on Vertex AI, find our documentation or reach out to your Google Cloud account representative to join the Trusted Tester Program."

This page might be somewhat helpful: https://cloud.google.com/vertex-ai/docs/generative-ai/image/... It also includes a link to the TTP form, although the form itself seems to make no reference to Imagen being part of the program anymore, confusingly. (Instead indicating that Imagen is GA.)

> GA.

Generally Available?

Re: Google Imagen 2

#14
post #10

To all the people saying “this sucks because we can’t use it” — there’s no real value in Google releasing this vs just making the announcement. This space is a race to the bottom, and there’s no significant profit being created in image gen right now (even if the service generates cashflow, the training and inference cost is insane). For the sake of team morale and legal risk, this announcement is totally enough, bet…

We can use it. It's generally available. We just can't find the page that explains how to use it or lets us test it.

Re: Google Imagen 2

#15

But how do we use it? Yet another documentation release by googling, promising impressive things that we cannot actually use, while the competition is readily available.

It says we can use it with their API. Would be good to have a link to it though.

Re: Google Imagen 2

#17
post #11
post #9

Earlier quoted context omitted.

> it's hard for me to tell I can only compare it to Stable Diffusion. But Imagen2 seems significant more advanced. Try to do anything with text and SDxl. It's not easy and often messes up. I don't think you can get a clean logo with multiple text areas on sdxl. Look at the prompt and image of the robin. That is mighty impressive.

Stability AI has gaps in SDXL for text, but they seem to do a better job with Deep Floyd ( https://github.com/deep-floyd/IF ). I have done a lot of interesting text things with Deep Floyd

Looks good. But 24GB of vram is quite a lot for 1024x1024

Re: Google Imagen 2

#18
post #13
post #7

Earlier quoted context omitted.

This page might be somewhat helpful: https://cloud.google.com/vertex-ai/docs/generative-ai/image/... It also includes a link to the TTP form, although the form itself seems to make no reference to Imagen being part of the program anymore, confusingly. (Instead indicating that Imagen is GA.)

> GA. Generally Available?

Yes
Post reply on HN