I think where computers fall short is in two areas:
1) The rate of errors hasn’t hit the inflection point of being comparable to day-to-day intra-human interaction, and
2) There is no good mechanism for detecting and correcting errors. At least, none that I’ve seen.
That second one is important. When I hear my wife ask me “Please, hand me a tractor”, I realize that I must’ve misheard, and ask her to clarify “what?” With speech recognition, I either have to manually re-read and modify recognized text, or cancel and repeat the entire request. Both take time, and negate some of the efficiencies of using speech recognition.