There are some obvious mistakes in tools like this. Such as: Human faces are wrong, writing is usually scrambled, fingers look weird etc... Do you know if we need to have a major breakthrough similar to what happened 6 months ago to fix this or could these be fixed with incremental improvements in current techniques / datasets?
I think a lot of this has been solved in DALL-E already. It's pretty good at right-looking faces, and fingers. Text not so much... but it does appear that whatever OpenAI are doing behind the scenes, that's getting better too.