I find it interesting that Microsoft had _so much of this_ 9 years go. I developed with the HoloLens 1 around 2016 and I recall:
* Fingers together gesture clicking.
* Voice activated menu navigation. It was glitchy/I never used it though. In 2023 these sorts of systems are much better though.
* No controllers. It was all gesture based. Opening the start menu required your hand upturned, fingers together then outstretch. Kind of like an "open" gesture.
* "Pointing" based on head looking vector, which was annoying.
* Spatial anchors and being able to remember past spaces and how you used them. There was a whole set of SDK APIs for the spatial stuff built into Win 10
The Vision Pro iterates on some of these. Eye tracking for pointing and more cameras for tracking hand pose for clicking, to reduce the annoyance/strain. IMO all mandatory for long term usage (I frequently got motion sick with the HoloLens and the 3 months of debugging/developing on it were a challenge).
Maybe if Microsoft had switched to more commonplace display technology and continued to iterate on their product instead of let it languish, they could have had a solid competitor to this, if not maybe be first to the market.