Good FLOSS speech recognition and TTS is badly needed. Such interaction should not be left to an oligoply with bad history of not respecting users freedoms and privacy.
Example: https://www.youtube.com/watch?v=tfcme7maygw
Granted, maybe this is "not good enough", but I feel like I got pretty far with pico2wave, pocketsphinx plus 1980's Zork level "comprehension" technology.
And the open source status of pico2wave is a bit questionable, I'll grant you that.
A bit more detail about the implementation here: https://scaryreasoner.wordpress.com/2016/05/14/speech-recogn...