Live data from Hacker News

Show HN: A Dalle-3 and GPT4-Vision feedback loop

dalle.party

131–140 of 156 posts

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#131
post #91

Earlier quoted context omitted.

The thing that is truly mindboggling to me is that THE SHADOWS IN THE IMAGES ARE CORRECT. How is that possible??? Does DALL-E actually have a shadow-tracing component?

Research into the internals of the networks have shown that they figure out the correct 2.5D representation of the scene before the RGB textures (internally), so yes it seems they have an internal representation of the scene and therefore can do enough inference from that to make shadows and light seem natural. I guess it's not that far-fetched as your brain has to do the same to figure out if a scene (or an AI-gener…

What does 2.5D mean?

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#133
post #131

Earlier quoted context omitted.

Research into the internals of the networks have shown that they figure out the correct 2.5D representation of the scene before the RGB textures (internally), so yes it seems they have an internal representation of the scene and therefore can do enough inference from that to make shadows and light seem natural. I guess it's not that far-fetched as your brain has to do the same to figure out if a scene (or an AI-gener…

What does 2.5D mean?

It means you should be worried about the guy she told you not to worry about

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#134
Nice! I prototyped a manual version of this a while ago. https://twitter.com/conradgodfrey/status/1712564282167300226

I think the thing that strikes me is that the default for chatGPT and the API is to create images in "vivid" mode. There's some interesting discussion on the differences between the "vivid" and "natural" here https://cookbook.openai.com/articles/what_is_new_with_dalle_...

I think these contribute to the images becoming more surreal - would be interested to compare to natural mode - it looks like you're using vivid mode based on the examples?

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#136
I purposely gave it some weird instructions to show the progress of the universe from the Big Bang to present day Earth. It showed the 8 stages from my prompt in each image and started to iterate over it, and then on image four I got a 400 error: Error: 400 Your request was rejected as a result of our safety system. Your prompt may contain text that is not allowed by our safety system. Interesting.

https://dalle.party/?party=EdpKnnBC

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#137

The "create text version of image" prompt matters a ton. I tried three, demo here: default https://dalle.party/?party=JfiwmJra hyper-long + max detail + compression - This shows that with enough text, it can do a really good job of reproducing very, very similar images https://dalle.party/?party=QtEqq4Mu hyper-long + max detail + compression + telling it to cut all that down to 12 words - This seems okay. I might be…

    4
    GPT4 vision prompt generated from the previous image:
    I'm sorry, I cannot assist with this request.
Is that because it's gradually made the spaceship look more like some sort of RPG backpack, so now it thinks it's being asked to describe prompts to create images of weaponry and that's deemed unsafe?

Re: Show HN: A Dalle-3 and GPT4-Vision feedback loop

#140
post #22

Cool idea! I made one with the starting prompt "an artificial intelligence painting a picture of itself": https://dalle.party/?party=wszvbrOx It consistently shows a robot painting on a canvas. The first 4 are paintings of robots, the next 3 are galaxies, and the final 2 are landscapes.

In a few these pictures it seems to be heavily influenced by the adaptation of I Robot with Will Smith in it for what robots look like.
Post reply on HN