The big hurdle for voice UI is discoverability.
That's probably why it's most successful in a handful of verticals, like timers and music control. There are a finite list of likely commands, and people can probably guess them first time.
When you move much beyond that, you start getting a high miss rate because there's no good way to pull a list of valid commands or options.
Importantly, especially for home automation, "option guessability" is a big thing-- you might know the commands, but the relevant options are specific to each location, and you were likely not there when it was configured. How do you know that the light fixtures are named "Left sconce, right sconce, ceiling fan, and table lamp", and not "north, south, Bob, and Matilda?"
Voice UI is also a terrible situation for error recovery. I can't figure out a voice-only UI that would handle a simple case like "which light do you want to turn on?" well. At best, it could read you a list of all the devices that might be relevant for the question, but that's so low information-density and will feel like a bad telephone IVR service.
The "Siri with a screen" design convention might have been sensible, as that would have provided a way to disambiguate queries quickly, with buttons to show the options, and at idle, a browsable list of supported commands.
Of course, at that point, you're making a more obvious machine with a (voice-based) command-line. That's shatters the illusion they're selling of this being your secretary/concierge, with arbitrary and broad capabilities, who happens to be a robot.