For solving long term tasks like finding things that aren't there, you can turn the annotated scene into a templated description and feed it to a large-enough model trained on interactive fiction. You are standing in a kitchen. Ahead of you to your right there is a large refrigerator with the handle on the right side. There is a set of cabinets to your left with a plate sitting on the counter above them. > get beer Y…
Great, now we can teach robots to wander around rooms looking for things, saying "keys, keys, keys... where would I put keys?"
You could take video data and have fuzzy identification of objects moving around, then throw away the video and keep track of the objects, the blue floppy thing (gloves) and the metal shiny deforming things (keys) then have a more constructive dialog about the keys. A voice responding, what do the keys look like? Is there a blue square thing on the key ring? The less identifiable the object the funnier the discussion. What shirt? You have many shirts! Oh, the blue one, you have 4 of those, one in the sink, one behind the bed, one in the laundry basket, one in the closet. Oh the one with stripes! Why didn't you say so, it's behind the bed bro.
It could also ask you if they are suppose to be on the outside in the front door after you close it.