Earlier quoted context omitted.
I’ve tried it on simple UI tasks. Give it screenshot to work towards (or figma MCP), let it get screenshots from chrome to check its work. The existing models are surprisingly bad at it.
It's really difficult to understand what your definition of surprisingly bad is. What was it continuously having problems with?
Or guessing colors rather than sampling from the image or pulling from figma is another stupid thing they do constantly.