I figured this would quickly go off the rails into surreal territory, but instead it ended up being progressive technological de-evolution. Starting prompt: "A futuristic hybrid of a steam engine train and a DaVinci flying machine" Results: https://dalle.party/?party=14ESewbz (Addendum: In case anyone was curious how costs scale by iteration, the full ten iterations in this result billed $0.21 against my credit balan…
Here's a second run of the same starting prompt, this time using the "make it more whimsical" modifier. It makes a difference and I find it fascinating what parts of the prompt/image gain prominence during the evolutions. Starting prompt: "A futuristic hybrid of a steam engine train and a DaVinci flying machine" Results: https://dalle.party/?party=qLHPB2-o Cost: Eight iterations @ $0.44 -- which suggests to me that t…
Show HN: A Dalle-3 and GPT4-Vision feedback loop
141–150 of 156 posts
Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop
#142It's pretty fun to mess with the prompt and see what you can make happen over the series of images. Inspired by a recent Twitter post[1], I set this one up to increase the "intensity" each time it prompted. The starting prompt (or at least, the theme) was suggested by one of my kids. Watch in awe as a regular goat rampage accelerates into full cosmic horror universe ending madness. Friggin awesome : https://dalle.par…
Great idea asking it to increase the intensity each run. This made my evening!
> Write a prompt for an AI to make this image. Just return the prompt, don't say anything else, but also, increase the intensity of any adjectives, resulting in progressively more fantastical and wild prompts. Really oversell the intensity factor, and feel free to add extra elements to the existing image to amp it up.
I played with it a bit before I got results I liked - one of the key factors, I think, was giving the model permission to add stuff to the image, which introduced enough variation between images to have a nice sense of progression. Earlier attempts without that instruction were still cool, but what I noticed was that once you ask it to intensify every adjective, you pretty much go to 11 within the first iteration or two - so you wind up having 1 image of a silly cat or goat and then 7 more images of world-shattering kaiju.
The goat one (which again, was an idea from one of my kids) was by far the best in terms of "progression to insanity" that I got out of the model. Really fun stuff!
Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop
#143Here's a custom prompt that I enjoyed: "Think hard about every single detail of the image, conceptualize it including the style, colors, and lighting. Final step, condensing this into a single paragraph: Very carefully, condense your thoughts using the most prominent features and extremely precise language into a single paragraph." https://dalle.party/?party=1lSMniUP https://dalle.party/?party=cEUyjzch https://dalle.…
The thing that is truly mindboggling to me is that THE SHADOWS IN THE IMAGES ARE CORRECT. How is that possible??? Does DALL-E actually have a shadow-tracing component?
Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop
#144Earlier quoted context omitted.
The entire thing is frontend only (except for the share feature) so the server never sees your key. You can validate that by watching the network tab in developer console. You can also make a new / revoke an API key to be extra sure.
Please make a new API key folks. There's a lot of tricks to scrape a text box and watching the network tab isn't enough for safety.
Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop
#145Earlier quoted context omitted.
Research into the internals of the networks have shown that they figure out the correct 2.5D representation of the scene before the RGB textures (internally), so yes it seems they have an internal representation of the scene and therefore can do enough inference from that to make shadows and light seem natural. I guess it's not that far-fetched as your brain has to do the same to figure out if a scene (or an AI-gener…
What does 2.5D mean?
Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop
#146Earlier quoted context omitted.
Research into the internals of the networks have shown that they figure out the correct 2.5D representation of the scene before the RGB textures (internally), so yes it seems they have an internal representation of the scene and therefore can do enough inference from that to make shadows and light seem natural. I guess it's not that far-fetched as your brain has to do the same to figure out if a scene (or an AI-gener…
Interesting! Do you have a link to that research?
It's a very interesting paper.
"Even when trained purely on images without explicit depth information, they typically output coherent pictures of 3D scenes. In this work, we investigate a basic interpretability question: does an LDM create and use an internal representation of simple scene geometry? Using linear probes, we find evidence that the internal activations of the LDM encode linear representations of both 3D depth data and a salient-object / background distinction. These representations appear surprisingly early in the denoising process−well before a human can easily make sense of the noisy images."
Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop
#147Earlier quoted context omitted.
I'd be interested to see how much of this is because the model doesn't know what it's looking at and how much is because describing picture with a short amount of text is inherently very lossy. Maybe one way to check would be doing this with people. Get 8 artists and 7 interpreters, craft the initial message, and compare the generational differences between the two sets?
Example: https://dalle.party/?party=42riPROf > Create an image of an anthropomorphic orange tabby cat standing upright in a kung fu pose, surrounded by a dozen tiny elephants wearing mouse costumes with mini trumpets, all gazing up in awe at a gigantic wheel of Swiss cheese that hovers ominously in the background. That's hilarious, but also hilariously wrong on almost every detail. There's a huge asymmetry in apparen…
Edit:In this case, it appears that it was a vanilla prompt "Write a prompt for an AI to make this image. Just return the prompt, don't say anything else.'
Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop
#148It's pretty fun to mess with the prompt and see what you can make happen over the series of images. Inspired by a recent Twitter post[1], I set this one up to increase the "intensity" each time it prompted. The starting prompt (or at least, the theme) was suggested by one of my kids. Watch in awe as a regular goat rampage accelerates into full cosmic horror universe ending madness. Friggin awesome : https://dalle.par…
Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop
#149Here's a custom prompt that I enjoyed: "Think hard about every single detail of the image, conceptualize it including the style, colors, and lighting. Final step, condensing this into a single paragraph: Very carefully, condense your thoughts using the most prominent features and extremely precise language into a single paragraph." https://dalle.party/?party=1lSMniUP https://dalle.party/?party=cEUyjzch https://dalle.…
It's very interesting to observe how the relationship between the wolf and Redhood evolved from dark and menacing to serene and friendly.
Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop
#150This reminds me of the party game Telestrations where players go back and forth between drawing and writing what they see. It's hilarious to see the result because you anticipate what the next drawing will be while reading the prompt. I'd love to see an alternative viewing mode here which shows the image and the following prompt. Then you need to click a button to reveal the next image. This allows you to picture in…