Background: researched this space for a graduate degree.
There are a few issues that are unanswered by this video (which isn't intended to be a technical deep dive, but I don't see any related links in the video description):
1. How do these glasses handle multiple simultaneous speakers? Based on the display I saw, it shows the speakers' words sequentially, which starts to fall apart in real-world environments, especially group conversations. This is a big problem, and wider adoption is contingent on handling this elegantly.
2. These appear to be the classic "smart glasses" display style that's pervasive in consumer head-worn displays today, where content is projected at a fixed depth in front of the wearer. Because the captions aren't anchored at the same focal distance as the speaker, the wearer's eyes will swap between the captions and the speaker's faces, which is a tiring activity, and can make the wearer feel like they're not part of the conversation or being rude.
3. As mentioned by another commenter, this is a useful idea for people who lose their hearing later in life. That said, this is less (although certainly still) useful for people who have congenital hearing loss and primarily communicate via ASL.
All in all, it's exciting to see growing interest in this space, as it's easily extendable to people learning a new language or navigating a foreign country. I think offloading the speech-to-text to a tethered mobile device is a good choice (though it would be nice to do low-latency wireless transmission).