Why is this text-based image research advancing so rapidly? Is there a market application they're aiming for? Seems like multiple teams have been on multiple models and I see a new one every week.
In terms of theory, such systems are candidates for a generic perception engine you might use in say, a robot with cameras, speaker and a microphone.
Perception is just one aspect of intelligence, but this research ultimately makes it possible for a machine to encode data semantically.