The language of high-level art-direction can be way more complex than one might assume. I wonder how this model might cope with the following: ‘Decrease high-frequency features of background.’ ‘Increase intra-contrast of middle ground to foreground.’ ‘Increase global saturation contrast.’ ‘Increase hue spread of greens.’
Hopefully in a couple of years when things have matured more there will be more models capable of handling said requests
The most precise models are actually anime models because the users have got high standards for telling the machine what they expect of it and the databases are quite well annotated (booru tags)