Earlier quoted context omitted.
> "[...]the final result is more cohesive around a single them than the original idea." That's an observation worth investigating. Here's another set of data points to see if there's more to it... Input prompt: "Six robots on a boat with harpoons, battling sharks with lasers strapped to their heads" GPT4V prompt: "Write a prompt for an AI to make this image. Just return the prompt, don't say anything else. Make it fu…
Both of your examples seem to start with two subjects (steam engine/flying machine and shark/robot), and throughout the animation one of them gets more prominence until the other is eventually dropped altogether.
GPT4V instructions for all tests: "Write a prompt for an AI to make this image. Just return the prompt, don't say anything else. Make it weirder."
From what you'll see in the results there's possible evidence of bias towards the first subject listed in a prompt, making it the object of fixation through the subsequent iterations. I'll also speculate that "gnomes" (and their derivations) and "cosmic images" are over-represented as subjects in the underlying training data. But that's wild speculation based on an extremely small sample of results.
In any case, playing around with this tool has been enjoyable and a fun use of API credits. Thank you @z991 for putting this together and sharing it!
------ Test 1 ------
Prompt: "Two garden gnomes, a sentient mushroom, and a sugar skull who once played a gig at CBGB in New York City converse about the boundaries of artificial intelligence."
Result: https://dalle.party/?party=ZSOHsnZe
------ Test 2 ------
Prompt: "A sentient mushroom, a sugar skull who once played a gig at CBGB in New York City, and two garden gnomes converse about the boundaries of artificial intelligence."
Result: https://dalle.party/?party=pojziwkU
------ Test 3 ------
Prompt: "A sugar skull who once played a gig at CBGB in New York City, a sentient mushroom, and two garden gnomes converse about the boundaries of artificial intelligence."