Earlier quoted context omitted.
Can you really distinguish sound emitted from the speaker from someone with a hoarse voice? Furthermore, what about medium? Traveling different medium should be considered.
Well, the device knows exactly what signal it's putting out through the speaker, so it can predict what the microphone will pick up. It doesn't know what someone with a hoarse voice is about to say.
And further, the attack described is a sentence that doesn't sound to a human like "Ok Google" or "Hey Siri", or whatever