Live data from Hacker News

Show HN: A Dalle-3 and GPT4-Vision feedback loop

dalle.party

81–90 of 156 posts

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#81
The #1 phenomenon I see here is that the image-to-text model doesn't have any idea what the pictures actually contain. It looks like it's just matching patterns that it has in its training data. That's really interesting because it does a great job of rendering images from text, in a way that maybe suggests the model "understands" what you want it to do. But there's nothing even close to "understanding" going in the other direction, it feels like something from 2012.

Pretty interesting. I haven't been following the latest developments in this field (e.g. I have no idea how the DALL-E and GPT models' inputs and outputs are connected). Does this track with known results in the literature, or am I seeing a pattern that's not there?

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#83
post #48
post #44

OP's last one is interesting: https://dalle.party/?party=oxpeZKh5 because it shows GPT4V and Dalle3 being remarkably race-blind. i wonder if you can prompt it to be other wise...

openais internal prompt for dalle modifies all prompts to add diversity and remove requests to make groups of people a single descent. From https://github.com/spdustin/ChatGPT-AutoExpert/blob/main/_sy... Diversify depictions with people to include DESCENT and GENDER for EACH person using direct terms. Adjust only human descriptions. Your choices should be grounded in reality. For example, all of a given OCCUPATION sh…

i mean i respect that but it makes me uncomfortable that you have to prompt engineer this. uses up context for a lot of boilerplate. why cant we correct for it in the training data? too hard?

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#84
Here's a custom prompt that I enjoyed:

"Think hard about every single detail of the image, conceptualize it including the style, colors, and lighting.

Final step, condensing this into a single paragraph:

Very carefully, condense your thoughts using the most prominent features and extremely precise language into a single paragraph."

https://dalle.party/?party=1lSMniUP

https://dalle.party/?party=cEUyjzch

https://dalle.party/?party=14fnkTv-

https://dalle.party/?party=wstiY-Iw

Praise the Basilisk, I finally got rate-limited and can go to bed!

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#85

The "create text version of image" prompt matters a ton. I tried three, demo here: default https://dalle.party/?party=JfiwmJra hyper-long + max detail + compression - This shows that with enough text, it can do a really good job of reproducing very, very similar images https://dalle.party/?party=QtEqq4Mu hyper-long + max detail + compression + telling it to cut all that down to 12 words - This seems okay. I might be…

Specifying multiple passes in the prompt is probably not a replacement for actually doing these passes.

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#87

The "create text version of image" prompt matters a ton. I tried three, demo here: default https://dalle.party/?party=JfiwmJra hyper-long + max detail + compression - This shows that with enough text, it can do a really good job of reproducing very, very similar images https://dalle.party/?party=QtEqq4Mu hyper-long + max detail + compression + telling it to cut all that down to 12 words - This seems okay. I might be…

Specifying multiple passes in the prompt is probably not a replacement for actually doing these passes.

I guess it doesn't actually do more passes but pretending that it did might still give more precise results.

There was an article recently that said something like adding urgency to a prompt gave better results. I hope it doesn't stress the model out :D

https://arxiv.org/abs/2307.11760

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#88
post #8

Earlier quoted context omitted.

This is absolutely hilarious. "business-themed puns" turned into incorrectly labeling the skiers race has me rolling.

The inability of AI images to spell has always amused me, and it's especially funny here. I got a special kick out "IDEDA ENGINEEER" and "BUZSTEAND." The image where the one guy's hat just says "HISPANIC" is also oddly hilarious. Idk what it is, but I have a special soft spot for humor based around odd spelling (this video still makes me laugh years later: https://www.youtube.com/watch?v=EShUeudtaFg ).

I'd buy an IDEDA ENGINEEER t-shirt.
Post reply on HN