Can do UI to frontend. Seems to understand the UI graphical elements and layout, not just text https://twitter.com/skirano/status/1706823089487491469
Can describe comic images accurately, panel by panel - https://twitter.com/ComicSociety/status/1698694653845848544?...
Lots of examples here also - https://www.reddit.com/r/ChatGPT/comments/16sdac1/i_just_got...
It's Computer Vision on Steroids basically.
Multi-modality is pretty low hanging fruit so i'm glad we're finally getting started on that. Imagine if GPT-4 could manipulate sound and images even half as well as it could manipulate text. We still don't have a large scale multi-modal model trained from scratch so a lot of possible synergistic effects are still unknown.