I noticed that your profile mentioned "recording engineer" so maybe some concrete numbers related to digital audio technology will put boundaries on plausible scenarios.
We assume either of 2 engineering designs:
(#1) the trigger word "Alexa" is detected within an embedded chip. The DSP (digital signal processing) intelligence for analyzing sound waveforms is inside the device. Therefore, the words spoken after "Alexa" are then sent to the cloud.
(#2) the trigger word "Alexa" (and/or other words) are detected remotely via cloud computers. There is no "smart" DSP chip within the Echo device. That means that the device must send a constant 24/7 stream of digital waveforms to the cloud.
If we continue on the #2 scenario, we can guesstimate what data transfer volumes would look like. To be conservative, we use 8kHz 8-bit audio as the parameters which is telephone quality. (Reliable voice recognition probably requires inputs with greater audio fidelity e.g. 16-bit 32kHz but we'll keep the 8kHz-8bit as a possible lower bound.)
Using 8kHz-8bit, it means that the device would have to stream 691 megabytes a day which leads to 20.7 gigabytes a month. Likewise on the back end, the amazon infrastructure would have to scale up to constantly analyze millions of parallel 24/7 digital waveforms. The amazon datacenters would be burning up terawatts of electricity to ignore the 99.99% of digital waveforms that is not the word "Alexa".
So, are there any consumer devices out there surreptitiously uploading 691 megabytes of digital waveforms (or any data) every single day? Is it realistic that Amazon would engineer the product to work like this?
I have a router that has a fallback option to a cellular connection in case my cable is disrupted. I and others would hate to get a surprise bill from Verizon/AT&T for going over my 2GB/month transfer limit if the amazon device was designed via scenario #2.
EDIT TO ADD scenario #3:
(#3) there are unpublicized/secret list of words in addition to the documented "Alexa" within the embedded chip's "vocabulary". Such words might be "vacation" and "book" and depending on the subsequent words sent to the cloud, you'd see ads for suntan lotion or Stephen King novels on your next visit to amazon.com. The chip's vocabulary may also include listening for transient sounds like dog barks or sneezes. You'd then get ads for dog food and cold medicine. In this scenario, a constant digital waveform is not uploaded 24/7 but extra trigger keywords unknown to the consumer causes more data to be sent than he/she agreed to.