What's a large language model doing when it's not being queried? Am I correct that they only compute information when dealing with a prompt? If so, that seems like a fundamental flaw. An actual "thinking machine" would be constantly running computations on its accumulated experience in order to improve its future output.
Then you can find a way to include a described video feed and method of movement into the mix.