It’s easy enough to test, but I don’t really need to.
Alexa Guard has had this functionality for a while, and I’d expect the folks here at HN to be able to infer a few things from the support link and basic reasoning.
Support link https://support.ring.com/hc/en-us/articles/360028205592-Usin...
So:
1) if an event is detected, you can listen to a 10 second clip or drop in (2-way call) to listen in or look.
2) Echo devices have relatively small amounts of RAM
3) Echo devices aren’t constantly hammering WiFi connections
From this, one should be able to deduce that the wakeword engine detects events and streams clips to servers only in situations that match events and settings to support these features. Why? Because processing, transit, and storage aren’t free, and one can’t store data in RAM that isn’t there or transmit data over WiFi without the physical layer showing signs of it. Furthermore, Amazon hasn’t cracked the code on hyper-efficient GB into KB lossless compression only to squirrel it away only for use in voice assistants.
Take the number of Alexa devices sold and run the numbers for all of those devices sending audio data to AWS all the time. The costs would be astronomical. The same goes for Google (though not with AWS). They’re no doubt incorporating the detectors into their on-device models.