It seems to me that the main innovation here is the quality of the microphone array.
As a professional sound recordist, the #1 challenge of recording from a fixed point is that the ambient noise and reflections within the room rapidly swamp the original signal when you record from a point source. You can hear someone talk from the far side of a room in person very easily, because your brain constantly compensates for the acoustic environment it is currently in. But when you hear a recording made in a different acoustic environment (eg a scene in a movie) then your tolerance for background noise is far lower, because you become acutely aware that the acoustics are not responsive to positional adjustments - in much the same way that the image on a screen is limited to a plane.
So when recording sound for film or video, we tend to use special microphones with long barrels (which are highly directional) or fit actors and/or sets with very small microphones that only pick up sounds in close proximity and then transmit them by radio or wire. There are also parabolic microphones, but they're unwieldy and hard to focus plus they still pick up a lot of ambience, so they're better for things like sporting events where players repeatedly stand in predictable positions. The aim in recording sound this way is to get the actor's vocal performance with as little ambient noise as possible, which is then supplemented in post-production with additional recordings of background elements that can be layered in a controlled fashion. When recording on location rather on a sound stage, a large percentage of the takes are made for sound reasons; you would not believe how noisy the world is until you start trying to make quiet recordings of it. On almost every film project I have to have an argument with the producers at the early stage to be allowed (and paid) to come on location scouts, because most people are incapable of assessing the noise level of a location - their brains are so good at filtering out ambient noise and focusing on the conversations they're having about how the place looks that they are oblivious to how it sounds! I've been taken to what I was told was a quiet location only to discover that it was in the flight path of an airport 8-o
Anyway, the nice thing about this machine is the differential microphone array at the top. As well as providing a more accurate signal by simple differentiation, recording the device's own output and measuring what comes back in allows it to acoustically model the space it is in and then subtract that model from the input stream so as to isolate command spoken from across the room. I'd guess that most of this signal processing takes place on a DSP, and that the actual speech recognition is done in the cloud - though maybe not, as cheap CPUs pack so much punch nowadays. If you could hear the input to the speech recognition subsystem, it would sound oddly attenuated as it is stripped of any acoustic cues whatsoever.
I think the device will succeed or fail based on how semantically responsive it is - although different people will have different expectations and tolerances. For example:
You: Echo, I want to hear some new music!
Echo: How about the new album from XYZ?
You: Sure, I'll give that a try.
(music plays)
You: Echo, this music sucks.
(music keeps playing)
If Amazon (or anyone) can get a leg up on this sort of responsive conversation rather than just requiring the user to dictate commands all the time, they'll have a winner, even if it's little more than an Eliza front-end to a search engine.