Earlier quoted context omitted.
Anything that needs to overcome concepts which are disproportionately represented in the training data is going to give these models a hard time. Try generating: - A spider missing one leg - A 9-pointed star - A 5-leaf clover - A man with six fingers on his left hand and four fingers on his right You'll be lucky to get a 25% success rate. The last one is particularly ironic given how much work went into FIXING the ol…
It mostly depends on "how" the models work. Multi-modal unified text/image sequence to sequence models can do this pretty well, diffusion doesn't.
For example, to my knowledge ChatGPT is unified and I can guarantee it can't handle something like a 7-legged spider.