Live data from Hacker News

Imagen: An AI system that creates photorealistic images from input text

imagen.research.google

141–150 of 233 posts

Re: Imagen: An AI system that creates photorealistic images from input text

#141
post #130

Earlier quoted context omitted.

Which feels a bit sketchy(?) to me, seeing as all models are built on imagery scraped from the net without anyone's permission. It's one reason why I've spent the last month training my own models using my own imagery. If these stock photo sites had any brains, they would also start training models on images in their databases, especially since they already have everything sorted into categories based on keywords (wh…

I totally agree. Also, that sounds like a wonderful project :) You should post project updates somewhere to keep track of your model as it evolves!

Thanks! I have a g-doc where I've been documenting settings and progress here[0] although I just now realized I might not have been using my diffusion model correctly for most of my tests. Iteration 509 of my model I seem to have finally nailed it though! :) I partially "blame" Visions of Chaos since the amazing dev (or devs?) drops updates almost every day with new Machine Learning features, model training was only added recently. I must have reset something on accident.

Also I realize there's a lot of image prep work required, not to mention I have a less than ideal amount of VRAM (3060 Ti w/8GB but no monitors attached i.e. 8GB free) so I have to lower some settings. The source images have to be in 1:1 format (which none of my photos are) so I'm using a script to batch call ImageMagick's 'convert' to add white borders to the top/bottom, which results in my renders also having white borders.

[0] https://docs.google.com/document/d/1CnC5SaqpeJiQS-TlDS4trzJR...

Re: Imagen: An AI system that creates photorealistic images from input text

#142
post #74

Every single time I see an article about a new AI model that has a section called "societal impact" I know immediately they are not releasing the model, the training set, nothing... It seems to be the kind of bullshit statement that those companies put in place of "we paid $500k training this model and we're not giving it for free to anyone".

in the case of imagen, I suppose the cost is at least two orders of magnitude over 500k.

... no it definitely wasn't. that's $50m. read the paper, they tell you how long it took on a v4-256, which you know the public rental price for.

Re: Imagen: An AI system that creates photorealistic images from input text

#143

Earlier quoted context omitted.

Even if you don’t want to block generating porn, you still want to know if you’re getting it, because nobody wants to get porn when they didn’t ask for it. (and it may be illegal depending on your country) It’s easy to remove the filter from the SD scripts and that’s intentional.

Fair enough and props to Stability for giving people the benefit of the doubt, not to mention actually open sourcing almost all of their work, unlike other "open" projects. Speaking of which, Dall-E bans a decent chunk of the English dictionary in the hopes of preventing people from generating anything even remotely offensive, which can be somewhat frustrating.

someone told me bloody mary was banned from dalle

Re: Imagen: An AI system that creates photorealistic images from input text

#144
post #129

Earlier quoted context omitted.

> I make ai art all day long! I feel like this is the epitome of modern "content creation" Typing a few sentences into a software you barely understand, said software shits 15 jpegs out, 2 are good, "hey I make art". What a sad state of affair, tech is consuming everything and people are cheering, one more step on the path to being complete useless key pressers.

Photography is well known to have killed art. Press a button on a box you barely understand and the camera shits out a pic. "Hey I make art". What a sad state of affair, tech is consuming everything and people are cheering, one more step on the path to being complete useless key pressers.

Photography didn't replace drawing and painting.

Also if you are doing something interesting with digital photography it is definitely not just pressing a button.

Re: Imagen: An AI system that creates photorealistic images from input text

#145

Earlier quoted context omitted.

Pixel values are discrete (length x width x r256 x g256 x b256) and vertex values are continuous, so that is one major difference. Secondly, there's vastly more labeled image data in the world than 3D data, so creating a CLMP (contrastive language and mesh pairing) model is harder. It's very late but I may be able to give a much better answer on more of the nuances of 3D generation tomorrow.

Would voxels be easier than vertex-based meshes? I can imagine you'd have the problem of stray floating voxels then, which isn't as noticeable when it happens with 2D pixels.

The “hot new thing” is NeRF, neural radiance fields, which can take into account the way light interacts with the object (and hence you can correlate data from pictures taken at different angles)

Re: Imagen: An AI system that creates photorealistic images from input text

#146

Earlier quoted context omitted.

I’m not crazy about the tone you’re striking, but the bit about “software you barely understand” does ring true. Lots of folks who think hitting play in a notebook makes them some sort of AI-art-engineer these days. On the other hand, the whole point of automating creation is that you don’t need to understand the underlying mechanisms.

It's division of labour. The toolmaker has a different set of skills to the user of the tools.

Time spent designing a tool is time spent not using the tool to make art and honestly the types of brains that are good at designing art tools... often have very formulaic and structurally rigid ways of thinking about creating art (they tend to use tools for what they are intended for which almost no breakthrough in art has happened from). It is a good thing if artists can uses tools without having to understand the nitty gritty.

Re: Imagen: An AI system that creates photorealistic images from input text

#147

I'm surprised how very little this forum and the public in general knows or understands about the current capabilities and availability to these ai art generators. You have midjourney.com beta.dreamstudio.ai craiyon.com (real quick version no fuss, low quality) creator.nightcafe.studio and those are just some of the entry-level ones. I make ai art all day long! Check out my media feed on twitter @Sheilaaliens

> I make ai art all day long! I feel like this is the epitome of modern "content creation" Typing a few sentences into a software you barely understand, said software shits 15 jpegs out, 2 are good, "hey I make art". What a sad state of affair, tech is consuming everything and people are cheering, one more step on the path to being complete useless key pressers.

> software shits 15 jpegs out, 2 are good

Not really. This very much depends on the prompt. Just look at these, made these yesterday in a batch run. Same prompts, different seeds. This all what it made on seed it chose.

https://imgur.com/a/NGSvK48

Re: Imagen: An AI system that creates photorealistic images from input text

#148
post #143

Earlier quoted context omitted.

Fair enough and props to Stability for giving people the benefit of the doubt, not to mention actually open sourcing almost all of their work, unlike other "open" projects. Speaking of which, Dall-E bans a decent chunk of the English dictionary in the hopes of preventing people from generating anything even remotely offensive, which can be somewhat frustrating.

someone told me bloody mary was banned from dalle

This list is already a month old but it should give you an idea of what kind of words are banned: https://www.reddit.com/r/dalle2/comments/wa3jt6/banned_words...

Re: Imagen: An AI system that creates photorealistic images from input text

#150
post #110
post #74

Every single time I see an article about a new AI model that has a section called "societal impact" I know immediately they are not releasing the model, the training set, nothing... It seems to be the kind of bullshit statement that those companies put in place of "we paid $500k training this model and we're not giving it for free to anyone".

I for one am glad that, for once in my life, an obviously huge advancement is taking into account the human impact of releasing the technology responsibly. Maybe it’s just a “BS statement” but given the major strides they’re making at removing racial/gender biases[1] from similar projects, I don’t think it’s just hot air. Especially given the phenomenon of “bias amplification”. Maybe I’m being too optimistic but eith…

>the major strides they’re making at removing racial/gender biases

They're literally just appending a race/gender string at the end [1]. In what world is that not just hot air?

[1] https://twitter.com/jd_pressman/status/1549523790060605440

Post reply on HN