Live data from Hacker News

Imagen: An AI system that creates photorealistic images from input text

imagen.research.google

31–40 of 233 posts

Re: Imagen: An AI system that creates photorealistic images from input text

#31

These tools are amazing for prototyping. I had an idea for a promotional poster, and seeing my idea just by writing it felt like magic. The generated image had too many artifacts to use, but gave me a guideline to follow when creating the real thing in Pixlr. AI content generation (text, image, source code, video, music) will be a huge boon for prototyping where applied judiciously.

Yep, I feel that exact way about Nvidia Canvas [1]. It does not produce anything even close to usable as a final product, but it can produce an amazing start to a concept.

[1] https://www.nvidia.com/en-us/studio/canvas/

Re: Imagen: An AI system that creates photorealistic images from input text

#32
post #12
post #2

Does anyone know if or when google intends to make this available for beta testing or public use?

Towards the bottom of the page they say: “The potential risks of misuse raise concerns regarding responsible open-sourcing of code and demos. At this time we have decided not to release code or a public demo. In future work we will explore a framework for responsible externalization that balances the value of external auditing with the risks of unrestricted open-access.”

Are they afraid of lawsuits or are they painfully regressive, prude moralists?

Stable diffusion was un-neutered within 24 hours of its public release and the worst people do with it is Emma Watson porn.

Re: Imagen: An AI system that creates photorealistic images from input text

#33

i only have a passing curiosity in these projects personally. can someone in the field explain why this has exploded recently? there seems to be a lot of these tools released recently (text to image) was there a major breakthrough? a new idea that pushed everyone forward? a recent sharing of talent between groups? edit: just another thought, are they just being posted to HN now, i don't see a date on the page for whe…

The previous models were either 1. Limited in their capacity to create something that looked very cool, or 2. Gigantic models that needed clusters of GPUs and lots of infrastructure to generate a single image.

One major thing that happened recently (2ish weeks ago) was the release of an algorithm (with weights) called stable diffusion, which runs on consumer grade hardware and requires about 8GB of GPU RAM to generate something that looks cool. This has opened up usage of these models for a lot of people.

example outputs with prompts for the curious: https://lexica.art/

Re: Imagen: An AI system that creates photorealistic images from input text

#34

Is Google going to share or do something other than better advertising with this? What do they plan to do with it? I'm frankly tired of these show-off blog posts by Google that neither make it into general hands nor are used for anything positive. This is just a link to their previous posted site btw, nothing new.

My guess is they're they'll use Imagen and LaMDA to build a "conversational" search experience of some kind. So, instead of providing a list of websites to go to, they'll synthesize an answer to the search query, with imagery to go along with it, and so on.

I don't know about the market at large, but I do not want that. I want the a search engine to just have a huge database of websites, and look up stuff in that database based on my query, spitting out a link to the page that matches best, then the one that matches second best etc etc.

Using ML to determine the order of matches is absolutely fine, but to "digest" the internet and cook up an answer to what I'm looking for without proper sourcing? I do not want that. I don't want to try and guess what biases the language model might have. It's way easier for me to gauge the bias of another human, and for that, I need to be sent to a page a human has written.

(Of course, I realize that the "blog written by a bot" genre of writing is also becoming more convincing, making this whole thing harder...)

Re: Imagen: An AI system that creates photorealistic images from input text

#35

i only have a passing curiosity in these projects personally. can someone in the field explain why this has exploded recently? there seems to be a lot of these tools released recently (text to image) was there a major breakthrough? a new idea that pushed everyone forward? a recent sharing of talent between groups? edit: just another thought, are they just being posted to HN now, i don't see a date on the page for whe…

Like anyone deeply in a field I know maybe several thousand people who could probably give a better answer, but I figure I'll give an effort to provide one since I don't see any good ones posted yet. The moment everyone knew this was going to be big was in 2019 when StyleGAN came out. They used a lot of tricks like aligning face features (like eyes) and had all their pictures of a single domain (the most famous being…

Thanks for answering. Since you mentioned your work on text-to-3d, what are the ways to enhance the image/3d model to actually be photo-(or rather reality)-realistic? Even (presumably) hand-picked examples from google on the linked page lack support bars of the sunglasses, include floating cups of wine with base-less Eiffel tower in the background.

P.S. It seems raccoons are unimaginable (even for AI) with any sunglasses: if photo-realistic mode is selected for a raccoon, changing to "wearing a sunglasses and" makes no difference :)

Re: Imagen: An AI system that creates photorealistic images from input text

#36
post #18

Earlier quoted context omitted.

Someone on the midjourney discord is making a comic with images generated by the bot. Another person is making a Magic The Gathering card pack with generated art (it looks good too!). People are already using it for stock photos One murky area we're still far away from but I'm curious to follow the developments on: AI-generated movies. Once generated clips gets good, what if some movie buff can just generate a movie,…

Lot of potential for movies- using AI to up res, AI to turn a 2d movie 3d, more advanced would be new or edited scenes. Or how about translating a live action movie to a cartoon and vice versa. Or a different style or tone. Run The Lord of the Rings through a cyberpunk filter.

I'm eager for someone to start recreating the missing episodes of Doctor Who. The audio still exists and there are publicity shots from many of the missing episodes

Re: Imagen: An AI system that creates photorealistic images from input text

#37
post #11

Smart for Dalle, Midjourney, and Stable Diffusion to capitalize on this quickly. It looks like the technology is being commoditized at rapid speed. I wonder what’s next.

People are getting into "prompt-craft". People have also started pasting different AI tools together like putting these images into CogVideo[0] to generate actual videos (takes ages and usually the length of a gif currently). Others have realized you can run these images through GANs to make faces look extremely convincing I think "what's next" is fitting these tools together in a larger system [0] https://replicate.…

Is there a good central source that lists out all of the currently available tools?

Re: Imagen: An AI system that creates photorealistic images from input text

#38
post #3

Not a single human face in these samples. I wonder how well it does on faces.

Saul Goodman if he was a character in Twin Peaks:

https://media.discordapp.net/attachments/999426920376717513/...

Generated with Midjourney Beta

Re: Imagen: An AI system that creates photorealistic images from input text

#39

These tools are amazing for prototyping. I had an idea for a promotional poster, and seeing my idea just by writing it felt like magic. The generated image had too many artifacts to use, but gave me a guideline to follow when creating the real thing in Pixlr. AI content generation (text, image, source code, video, music) will be a huge boon for prototyping where applied judiciously.

> Creating the real thing in Pixlr

is it that good now? note that Pixlr was bought by Google

Re: Imagen: An AI system that creates photorealistic images from input text

#40

i only have a passing curiosity in these projects personally. can someone in the field explain why this has exploded recently? there seems to be a lot of these tools released recently (text to image) was there a major breakthrough? a new idea that pushed everyone forward? a recent sharing of talent between groups? edit: just another thought, are they just being posted to HN now, i don't see a date on the page for whe…

The previous models were either 1. Limited in their capacity to create something that looked very cool, or 2. Gigantic models that needed clusters of GPUs and lots of infrastructure to generate a single image. One major thing that happened recently (2ish weeks ago) was the release of an algorithm (with weights) called stable diffusion, which runs on consumer grade hardware and requires about 8GB of GPU RAM to generat…

Is Lexica finding results previously computed? Or generating them? I could only work with very simple queries like "photo of a cat".
Post reply on HN