Live data from Hacker News

Show HN: A Dalle-3 and GPT4-Vision feedback loop

dalle.party

91–100 of 156 posts

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#91

Here's a custom prompt that I enjoyed: "Think hard about every single detail of the image, conceptualize it including the style, colors, and lighting. Final step, condensing this into a single paragraph: Very carefully, condense your thoughts using the most prominent features and extremely precise language into a single paragraph." https://dalle.party/?party=1lSMniUP https://dalle.party/?party=cEUyjzch https://dalle.…

The thing that is truly mindboggling to me is that THE SHADOWS IN THE IMAGES ARE CORRECT. How is that possible??? Does DALL-E actually have a shadow-tracing component?

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#92

Here's a custom prompt that I enjoyed: "Think hard about every single detail of the image, conceptualize it including the style, colors, and lighting. Final step, condensing this into a single paragraph: Very carefully, condense your thoughts using the most prominent features and extremely precise language into a single paragraph." https://dalle.party/?party=1lSMniUP https://dalle.party/?party=cEUyjzch https://dalle.…

> https://dalle.party/?party=14fnkTv-

Interesting that for one and only one iteration, the anthropomorphized cardboard boxes it draws are almost all Danbo: https://duckduckgo.com/?q=danbo+character&ia=images&iax=imag...

It was surprising to see a recognizable character in the middle of a bunch of more fantastical images.

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#93
post #8

Earlier quoted context omitted.

got really weird really fast https://dalle.party/?party=7cnx55yN

This is absolutely hilarious. "business-themed puns" turned into incorrectly labeling the skiers race has me rolling.

Honestly, I'm really confused by how it was able to keep the idea of "business-themed puns" through so much of it. I don't understand how it was able to keep understanding that those weird letters were supposed to be "business-themed puns."

I don't think any human looking at drawing #3, which includes "CUNNFACE," "VODLI-EAPPERCO," "NITH-EASTER," "WORD," "SOCEIL MEDIA," and "GAPTOROU" would have worked out, as GPT did, that those were "pun-filled business buzzwords."

Is the previous prompt leaking? That is, does the GPT have it in its context?

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#94

The #1 phenomenon I see here is that the image-to-text model doesn't have any idea what the pictures actually contain. It looks like it's just matching patterns that it has in its training data. That's really interesting because it does a great job of rendering images from text, in a way that maybe suggests the model "understands" what you want it to do. But there's nothing even close to "understanding" going in the…

I'd be interested to see how much of this is because the model doesn't know what it's looking at and how much is because describing picture with a short amount of text is inherently very lossy.

Maybe one way to check would be doing this with people. Get 8 artists and 7 interpreters, craft the initial message, and compare the generational differences between the two sets?

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#95
post #75

Don’t get the significance, anyone one of those guys images could have been prompted the first time

It's a fun way to get guided variations.

Maybe you don't know what you specifically want you just want stylized gnomes so you write "a gnome on a spotted mushroom smoking a pipe, psychedelic, colorful, Alice in Wonderland style" and by the end of it you get that massively long and stylized prompt.

Maybe you do know what you want but you don't want to come up with an elaborate prompt so you steer it in a particular direction like the cat example.

For the first one you can get similar effects by asking for variations but it seems like this has a very different drift to it. Fun, albeit expensive in comparison.

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#97

Interesting how similar this is to my family's favorite game: pictograph. 1. You start by describing a thing. 2. The next person draws a picture of it. 3. The next next person describes the picture. repeat steps 2 and 3 until everyone has either drawn or described the picture. You then compare the first and last description... and look over the pictures. One of the best ever was: Draw a penguin. The first picture was…

[deleted]

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#98
post #26

I figured this would quickly go off the rails into surreal territory, but instead it ended up being progressive technological de-evolution. Starting prompt: "A futuristic hybrid of a steam engine train and a DaVinci flying machine" Results: https://dalle.party/?party=14ESewbz (Addendum: In case anyone was curious how costs scale by iteration, the full ten iterations in this result billed $0.21 against my credit balan…

I like how in #9 the carriage is on fire, or at least steaming disproportionately.

These images are incredible but I often notice stuff like this and it kind of ruins it for me.

#3 & #4 are good too, when the tracks are smoking, but not the train.

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#99
It's pretty fun to mess with the prompt and see what you can make happen over the series of images. Inspired by a recent Twitter post[1], I set this one up to increase the "intensity" each time it prompted.

The starting prompt (or at least, the theme) was suggested by one of my kids. Watch in awe as a regular goat rampage accelerates into full cosmic horror universe ending madness. Friggin awesome:

https://dalle.party/?party=vCwYT8Em

[1]: https://x.com/venturetwins/status/1728956493024919604?s=20

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#100
post #93
post #8

Earlier quoted context omitted.

This is absolutely hilarious. "business-themed puns" turned into incorrectly labeling the skiers race has me rolling.

Honestly, I'm really confused by how it was able to keep the idea of "business-themed puns" through so much of it. I don't understand how it was able to keep understanding that those weird letters were supposed to be "business-themed puns." I don't think any human looking at drawing #3, which includes "CUNNFACE," "VODLI-EAPPERCO," "NITH-EASTER," "WORD," "SOCEIL MEDIA," and "GAPTOROU" would have worked out, as GPT did…

It's probably just finding non-intuitive extrema in its feature space or something...

the whole thing with the text in the images reminds me of this: https://arxiv.org/abs/2206.00169

and I found myself that dall-e sometimes even likes to add gibberish text unpromtedly, often with letters containing some garbled versions of words from the prompt, or related words

Post reply on HN