Live data from Hacker News

Show HN: A Dalle-3 and GPT4-Vision feedback loop

dalle.party

141–150 of 156 posts

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#141
post #31
post #26

I figured this would quickly go off the rails into surreal territory, but instead it ended up being progressive technological de-evolution. Starting prompt: "A futuristic hybrid of a steam engine train and a DaVinci flying machine" Results: https://dalle.party/?party=14ESewbz (Addendum: In case anyone was curious how costs scale by iteration, the full ten iterations in this result billed $0.21 against my credit balan…

Here's a second run of the same starting prompt, this time using the "make it more whimsical" modifier. It makes a difference and I find it fascinating what parts of the prompt/image gain prominence during the evolutions. Starting prompt: "A futuristic hybrid of a steam engine train and a DaVinci flying machine" Results: https://dalle.party/?party=qLHPB2-o Cost: Eight iterations @ $0.44 -- which suggests to me that t…

The second picture reminds me of Back to the Future III.

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#142
post #112

It's pretty fun to mess with the prompt and see what you can make happen over the series of images. Inspired by a recent Twitter post[1], I set this one up to increase the "intensity" each time it prompted. The starting prompt (or at least, the theme) was suggested by one of my kids. Watch in awe as a regular goat rampage accelerates into full cosmic horror universe ending madness. Friggin awesome : https://dalle.par…

Great idea asking it to increase the intensity each run. This made my evening!

Thanks! This was the custom prompt I used:

> Write a prompt for an AI to make this image. Just return the prompt, don't say anything else, but also, increase the intensity of any adjectives, resulting in progressively more fantastical and wild prompts. Really oversell the intensity factor, and feel free to add extra elements to the existing image to amp it up.

I played with it a bit before I got results I liked - one of the key factors, I think, was giving the model permission to add stuff to the image, which introduced enough variation between images to have a nice sense of progression. Earlier attempts without that instruction were still cool, but what I noticed was that once you ask it to intensify every adjective, you pretty much go to 11 within the first iteration or two - so you wind up having 1 image of a silly cat or goat and then 7 more images of world-shattering kaiju.

The goat one (which again, was an idea from one of my kids) was by far the best in terms of "progression to insanity" that I got out of the model. Really fun stuff!

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#143
post #91

Here's a custom prompt that I enjoyed: "Think hard about every single detail of the image, conceptualize it including the style, colors, and lighting. Final step, condensing this into a single paragraph: Very carefully, condense your thoughts using the most prominent features and extremely precise language into a single paragraph." https://dalle.party/?party=1lSMniUP https://dalle.party/?party=cEUyjzch https://dalle.…

The thing that is truly mindboggling to me is that THE SHADOWS IN THE IMAGES ARE CORRECT. How is that possible??? Does DALL-E actually have a shadow-tracing component?

I randomly checked a few links here and shadows were correct in 2 images out of a dozen... and any people tend to be horrifying in many

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#144
post #61
post #28

Earlier quoted context omitted.

The entire thing is frontend only (except for the share feature) so the server never sees your key. You can validate that by watching the network tab in developer console. You can also make a new / revoke an API key to be extra sure.

Please make a new API key folks. There's a lot of tricks to scrape a text box and watching the network tab isn't enough for safety.

Who could scrape the text box in this scenario?

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#145
post #131

Earlier quoted context omitted.

Research into the internals of the networks have shown that they figure out the correct 2.5D representation of the scene before the RGB textures (internally), so yes it seems they have an internal representation of the scene and therefore can do enough inference from that to make shadows and light seem natural. I guess it's not that far-fetched as your brain has to do the same to figure out if a scene (or an AI-gener…

What does 2.5D mean?

You usually say 2.5D when it's a 3D but only from a single vantage point with no info of the back-facing side of objects. Like the representation you get from a depth-sensor on a mobile phone, or when trying to extract depth from a single photo.

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#146

Earlier quoted context omitted.

Research into the internals of the networks have shown that they figure out the correct 2.5D representation of the scene before the RGB textures (internally), so yes it seems they have an internal representation of the scene and therefore can do enough inference from that to make shadows and light seem natural. I guess it's not that far-fetched as your brain has to do the same to figure out if a scene (or an AI-gener…

Interesting! Do you have a link to that research?

Certainly: https://arxiv.org/abs/2306.05720

It's a very interesting paper.

"Even when trained purely on images without explicit depth information, they typically output coherent pictures of 3D scenes. In this work, we investigate a basic interpretability question: does an LDM create and use an internal representation of simple scene geometry? Using linear probes, we find evidence that the internal activations of the LDM encode linear representations of both 3D depth data and a salient-object / background distinction. These representations appear surprisingly early in the denoising process−well before a human can easily make sense of the noisy images."

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#147

Earlier quoted context omitted.

I'd be interested to see how much of this is because the model doesn't know what it's looking at and how much is because describing picture with a short amount of text is inherently very lossy. Maybe one way to check would be doing this with people. Get 8 artists and 7 interpreters, craft the initial message, and compare the generational differences between the two sets?

Example: https://dalle.party/?party=42riPROf > Create an image of an anthropomorphic orange tabby cat standing upright in a kung fu pose, surrounded by a dozen tiny elephants wearing mouse costumes with mini trumpets, all gazing up in awe at a gigantic wheel of Swiss cheese that hovers ominously in the background. That's hilarious, but also hilariously wrong on almost every detail. There's a huge asymmetry in apparen…

It is hard to tell without knowing the actual instructions given to GPT for how to create a description. You would expect a big difference if GPT was asked to create a whimsical and imaginative description vs a literal description with attention to detail accuracy.

Edit:In this case, it appears that it was a vanilla prompt "Write a prompt for an AI to make this image. Just return the prompt, don't say anything else.'

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#148

It's pretty fun to mess with the prompt and see what you can make happen over the series of images. Inspired by a recent Twitter post[1], I set this one up to increase the "intensity" each time it prompted. The starting prompt (or at least, the theme) was suggested by one of my kids. Watch in awe as a regular goat rampage accelerates into full cosmic horror universe ending madness. Friggin awesome : https://dalle.par…

So your kid is also playing goat simulator? =D

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#149

Here's a custom prompt that I enjoyed: "Think hard about every single detail of the image, conceptualize it including the style, colors, and lighting. Final step, condensing this into a single paragraph: Very carefully, condense your thoughts using the most prominent features and extremely precise language into a single paragraph." https://dalle.party/?party=1lSMniUP https://dalle.party/?party=cEUyjzch https://dalle.…

> https://dalle.party/?party=1lSMniUP

It's very interesting to observe how the relationship between the wolf and Redhood evolved from dark and menacing to serene and friendly.

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#150
post #7

This reminds me of the party game Telestrations where players go back and forth between drawing and writing what they see. It's hilarious to see the result because you anticipate what the next drawing will be while reading the prompt. I'd love to see an alternative viewing mode here which shows the image and the following prompt. Then you need to click a button to reveal the next image. This allows you to picture in…

Reminds me of exquisite corpse, where folks take turns drawing a piece / writing a paragraph and can only see the most recent one (https://austinkleon.com/2020/07/02/exquisite-corpse/)
Post reply on HN