Live data from Hacker News

Show HN: A Dalle-3 and GPT4-Vision feedback loop

dalle.party

41–50 of 156 posts

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#42
One reason this is good is that the default gpt4-vision UI is so insanely bad and slow. This just lets you use your capacity faster.

Rate limits are really low by default - you can get hit by 5 img/min limits, or 100 RPD (requests per day) which I think is actually implemented as requests per hour.

This page has info on the rate limits: https://platform.openai.com/docs/guides/rate-limits/usage-ti...

Basically, you have to have paid X amount to get into a new usage cap. Rate limits for dalle3/images don't go up very fast but it can't hurt to get over the various hurdles (5$, 50$, 100$) as soon as possible for when limits come down. End of the month is coming soon. It looks like most of the "RPD" limits go away when you hit tier 2 (having paid at least 50$ historically via API to them).

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#43
post #12

it seems like if you create a shareable link, then add more images, you can't create a new link with the new images

Yeah, that's a bug, I'll try to fix it tonight!

thanks for this! Basically the default UI they provide at chat.openai is so bad, nearly anything you would do would be an improvement.

* not hide the prompt by default * not only show 6 lines of the prompt even after user clicks * not be insanely buggy re: ajax, reloading past convos etc * not disallow sharing of links to chats which contain images * not artificially delay display of images with the little spinner animation when the image is already known ready anyway. * not lie about reasons for failure * not hide details on what rate limit rules I broke and where to get more information

etc

Good luck, thanks!

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#45
post #31
post #26

I figured this would quickly go off the rails into surreal territory, but instead it ended up being progressive technological de-evolution. Starting prompt: "A futuristic hybrid of a steam engine train and a DaVinci flying machine" Results: https://dalle.party/?party=14ESewbz (Addendum: In case anyone was curious how costs scale by iteration, the full ten iterations in this result billed $0.21 against my credit balan…

Here's a second run of the same starting prompt, this time using the "make it more whimsical" modifier. It makes a difference and I find it fascinating what parts of the prompt/image gain prominence during the evolutions. Starting prompt: "A futuristic hybrid of a steam engine train and a DaVinci flying machine" Results: https://dalle.party/?party=qLHPB2-o Cost: Eight iterations @ $0.44 -- which suggests to me that t…

I find it somewhat fascinating that in both examples, the final result is more cohesive around a single them than the original idea.

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#46

The "create text version of image" prompt matters a ton. I tried three, demo here: default https://dalle.party/?party=JfiwmJra hyper-long + max detail + compression - This shows that with enough text, it can do a really good job of reproducing very, very similar images https://dalle.party/?party=QtEqq4Mu hyper-long + max detail + compression + telling it to cut all that down to 12 words - This seems okay. I might be…

[deleted]

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#47
post #12

Earlier quoted context omitted.

Yeah, that's a bug, I'll try to fix it tonight!

thanks for this! Basically the default UI they provide at chat.openai is so bad, nearly anything you would do would be an improvement. * not hide the prompt by default * not only show 6 lines of the prompt even after user clicks * not be insanely buggy re: ajax, reloading past convos etc * not disallow sharing of links to chats which contain images * not artificially delay display of images with the little spinner an…

the new fancy animation for images is SO annoying

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#48
post #44

OP's last one is interesting: https://dalle.party/?party=oxpeZKh5 because it shows GPT4V and Dalle3 being remarkably race-blind. i wonder if you can prompt it to be other wise...

openais internal prompt for dalle modifies all prompts to add diversity and remove requests to make groups of people a single descent. From https://github.com/spdustin/ChatGPT-AutoExpert/blob/main/_sy...

    Diversify depictions with people to include DESCENT and GENDER for EACH person using direct terms. Adjust only human descriptions.

    Your choices should be grounded in reality. For example, all of a given OCCUPATION should not be the same gender or race. Additionally, focus on creating diverse, inclusive, and exploratory scenes via the properties you choose during rewrites. Make choices that may be insightful or unique sometimes.

    Use all possible different DESCENTS with EQUAL probability. Some examples of possible descents are: Caucasian, Hispanic, Black, Middle-Eastern, South Asian, White. They should all have EQUAL probability.

    Do not use "various" or "diverse"

    Don't alter memes, fictional character origins, or unseen people. Maintain the original prompt's intent and prioritize quality.

    Do not create any imagery that would be offensive.

    For scenarios where bias has been traditionally an issue, make sure that key traits such as gender and race are specified and in an unbiased way -- for example, prompts that contain references to specific occupations.

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#49
post #8

Earlier quoted context omitted.

got really weird really fast https://dalle.party/?party=7cnx55yN

This is absolutely hilarious. "business-themed puns" turned into incorrectly labeling the skiers race has me rolling.

BIZ NESS

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#50
post #31

Earlier quoted context omitted.

Here's a second run of the same starting prompt, this time using the "make it more whimsical" modifier. It makes a difference and I find it fascinating what parts of the prompt/image gain prominence during the evolutions. Starting prompt: "A futuristic hybrid of a steam engine train and a DaVinci flying machine" Results: https://dalle.party/?party=qLHPB2-o Cost: Eight iterations @ $0.44 -- which suggests to me that t…

I find it somewhat fascinating that in both examples, the final result is more cohesive around a single them than the original idea.

> "[...]the final result is more cohesive around a single them than the original idea."

That's an observation worth investigating. Here's another set of data points to see if there's more to it...

Input prompt: "Six robots on a boat with harpoons, battling sharks with lasers strapped to their heads"

GPT4V prompt: "Write a prompt for an AI to make this image. Just return the prompt, don't say anything else. Make it funnier."

Result: https://dalle.party/?party=pfWGthli

Cost: Ten iterations @ $0.41

(Addendum: I'd forgotten to mention that I believe the cost differential is due to the token count of each of the prompts. The first case mentioned had less words passed through each of the prompts than the later attempts when I asked it to 'make it whimsical' or 'make it funnier'.)

Post reply on HN