I'm confused about why you need yolo and llava. Can't you simply use yolo without a multimodal LLM? What does that add? You can use yolo to detect and grab screen coordinates on its own, right?
- "person": "get gender and age of this person in 5 words or less",
- "car": "get body type and color of this car in 5 words or less".
So YOLO gives the bounding box and rough category, while llava describes the object in more details.