Earlier quoted context omitted.
I agree that the platforms don't currently allow them default access to microphones, but that's not how it used to be. I also agree that we would have likely heard about this (old) capability, due to the number of people who would have worked on it (unless they were supported by a TLA). However, your description of what it would need to work is woefully naïve. There were startup funded apps that matched entire record…
Shazam was bought by Apple several years ago, but identifying music is a different beast from picking out the phrase Hey Siri/Google/Alexa.
Those are both additional proof cases, but I wasn't talking about either Shazam (which does a lot of processing remotely) or "wake words" although I'm familiar with the accuracy and compute power requirements for those. Multi-keyword detection for advertising with >50% FAR/FRR is much easier than either of them. A conversation is likely to hear the word (or related ones) mentioned multiple time by multiple parties and the matcher can be tuned to the individual locally.