Live data from Hacker News

Imagen: An AI system that creates photorealistic images from input text

imagen.research.google

171–180 of 233 posts

Re: Imagen: An AI system that creates photorealistic images from input text

#171
post #129

Earlier quoted context omitted.

Photography is well known to have killed art. Press a button on a box you barely understand and the camera shits out a pic. "Hey I make art". What a sad state of affair, tech is consuming everything and people are cheering, one more step on the path to being complete useless key pressers.

You still have to decide what you take a pic of, that's the artistic part. You have to chose a subject, lights, &c. it's still incredibly more complex than "hey google draw me a horse with a tuxedo" You can tell that right away seeing how many camera users never produce anything of quality. A 6 years old can press the button of a camera but no 6 years old will produce meaningful work. In the case of Ai art tools you…

>You still have to decide what you take a pic of, that's the artistic part.

You still have to select the 2 AI jpegs out of the 15, that's the artistic part.

Re: Imagen: An AI system that creates photorealistic images from input text

#172

Earlier quoted context omitted.

The AI is moving up the "content ladder" generating text -> now images -> next videos Smart for Google to invest in this because their business relies on third party content to exist (blogs, webpages, youtube videos) If they can vertically integrate their business to create the content AND own the discovery algorithms then they officially win the internet

Generating videos and 3D models is _much_ more difficult than images. You can’t just train off videos from the internet in the same way, because they don’t have sufficient text labels to understand them like CLIP does.

I wonder if subtitles could be used, so rather than describing the video, you just write a script and it generates video for you. I'm certainly no expert, but it does seem like there's a lot more data there.

Re: Imagen: An AI system that creates photorealistic images from input text

#173
post #129

Earlier quoted context omitted.

Photography is well known to have killed art. Press a button on a box you barely understand and the camera shits out a pic. "Hey I make art". What a sad state of affair, tech is consuming everything and people are cheering, one more step on the path to being complete useless key pressers.

You still have to decide what you take a pic of, that's the artistic part. You have to chose a subject, lights, &c. it's still incredibly more complex than "hey google draw me a horse with a tuxedo" You can tell that right away seeing how many camera users never produce anything of quality. A 6 years old can press the button of a camera but no 6 years old will produce meaningful work. In the case of Ai art tools you…

You still have to pick styles, tweaks, lighting, subject. Then iterate, set composition, guide the results. You need to pick the right models, params and samplers to get the style you want just like choosing your film/camera.

> it's still incredibly more complex than "hey google draw me a horse with a tuxedo"

And just like taking a photo of a random horse in a field you'll mostly get a pretty bland result.

You can argue it's simpler to create art with, which is an odd complaint, but if you're not working to make something great you generally won't get it - just like photography is much simpler than painting as it's "just press a button".

Re: Imagen: An AI system that creates photorealistic images from input text

#174
post #129

Earlier quoted context omitted.

Photography is well known to have killed art. Press a button on a box you barely understand and the camera shits out a pic. "Hey I make art". What a sad state of affair, tech is consuming everything and people are cheering, one more step on the path to being complete useless key pressers.

Photography didn't replace drawing and painting. Also if you are doing something interesting with digital photography it is definitely not just pressing a button.

Exactly, this won't replace all other forms of art and if you're doing the "just press the button" equivalent you won't get things that are that great.

At worst the complaint seems to be "with care you can more easily create good art" which is a very odd complaint.

Re: Imagen: An AI system that creates photorealistic images from input text

#175

Is anyone aware of people doing the same for NSFW images? They can easily wipe out an entire industry. No models to pay and to check for legal ages, infinite possibilities: just write the pic you want, the massive body part you want, how many genitals are involved and boom. You have your image.

[0] Was front page recently.

[0] (OBVIOUSLY NSFW) https://news.ycombinator.com/item?id=32572770

Re: Imagen: An AI system that creates photorealistic images from input text

#176
post #24

i only have a passing curiosity in these projects personally. can someone in the field explain why this has exploded recently? there seems to be a lot of these tools released recently (text to image) was there a major breakthrough? a new idea that pushed everyone forward? a recent sharing of talent between groups? edit: just another thought, are they just being posted to HN now, i don't see a date on the page for whe…

It is now available to lay people by just typing into a website. Months ago it was rather „use this Jupyter notebook“. So people are now using it for more serious stuff. For example, here is an RPG designer using Midjourney for illustrations: https://www.bastionland.com/2022/07/primeval-bastionland-pla...

A coworker and I were playing with DALL-E 2 yesterday, and I pointed out that while I don't think any major RPGs are going to be moving away from artists anytime soon, the quality of Ashcan Editions just jumped way, way up.

Re: Imagen: An AI system that creates photorealistic images from input text

#177
post #129

Earlier quoted context omitted.

Photography is well known to have killed art. Press a button on a box you barely understand and the camera shits out a pic. "Hey I make art". What a sad state of affair, tech is consuming everything and people are cheering, one more step on the path to being complete useless key pressers.

You still have to decide what you take a pic of, that's the artistic part. You have to chose a subject, lights, &c. it's still incredibly more complex than "hey google draw me a horse with a tuxedo" You can tell that right away seeing how many camera users never produce anything of quality. A 6 years old can press the button of a camera but no 6 years old will produce meaningful work. In the case of Ai art tools you…

> "hey google draw me a horse with a tuxedo"

I think this is an oversimplification.

There was a recent write up by a guy who used DALL-E to create his logo for his open source project. What was clear from that writeup is that it is still a process for getting exactly the look that one is aiming to achieve. Even with an AI, there are different styles, decisions, choices, and visual representations that have to be made.

Your position that with photography, you have to "chose a subject, lights, &c" doesn't change with AI generated artwork; one still has to describe the subject, color scheme, visual style, composition details, etc. for the AI to generate the image. Except that instead of composing a scene with makeup, props, and subjects, you do it textually.

I'd say that in some respect, it is far more "creative" than photography because it removes physical and real-world constraints from the artist which would otherwise require knowledge of CGI and digital tools.

> A 6 years old can press the button of a camera but no 6 years old will produce meaningful work

This is also true of even painting. Even a 6 year old can grab a paintbrush and paint without producing meaningful work. So that does not change with AI generated artwork. Yes, a 6 year old can describe a scene to an AI that generates some image -- just as a 6 year old can pick up a brush and apply paint to a canvas, but the likeliness of a 6 year old presenting the seed/input that the AI needs to generate something unique and of visual interest/originality is low just as it is with a paintbrush.

Re: Imagen: An AI system that creates photorealistic images from input text

#178

Earlier quoted context omitted.

I don't quite remember whether it was first used in Vit paper[1], but it's a fairly straight forward idea. You take the patches of an image like they are words in a sentence, reduce the size of the patch(num_of_pixel x num_of_pixel) with a linear projection so that we can actually process it and get rid of sparse pixel information, add in positional encodings to put in location information of the patch and treat them…

> reduce the size of the patch(num_of_pixel x num_of_pixel) with a linear projection What does that mean? (Thanks for the explanation)

The flattened image patch of width and height PxP pixels gets multiplied with a learnable matrix of dimension P^2xD where D is the size of the patch embedding. In other words, it’s a linear transformation that reduces the dimensionality of the image patch.

Re: Imagen: An AI system that creates photorealistic images from input text

#179

Is anyone aware of people doing the same for NSFW images? They can easily wipe out an entire industry. No models to pay and to check for legal ages, infinite possibilities: just write the pic you want, the massive body part you want, how many genitals are involved and boom. You have your image.

I agree, it does seem like this makes a lot of sense for the adult industry. But seemingly all of the models that get released have a built in censor for that sort of thing from what I understand.

Re: Imagen: An AI system that creates photorealistic images from input text

#180

Earlier quoted context omitted.

The AI is moving up the "content ladder" generating text -> now images -> next videos Smart for Google to invest in this because their business relies on third party content to exist (blogs, webpages, youtube videos) If they can vertically integrate their business to create the content AND own the discovery algorithms then they officially win the internet

Generating videos and 3D models is _much_ more difficult than images. You can’t just train off videos from the internet in the same way, because they don’t have sufficient text labels to understand them like CLIP does.

Oh but they have sound which can be annotated much faster/more efficiently. You also potentially have screenplay but the amount of training data is probably too less and sparse.

FWIW, I don’t think the AI systems will generate a whole video by itself - it’ll be some form of image to image generation where an artist will render a rough sketch of the scene and the AI will fill in the details, frame by frame.

Post reply on HN