For me, the thing that killed voice commands had nothing to do with speech technology. It was latency and error handling. At the start of my morning commute, I would say, "Ok Google, navigate to work". Often, this would fail because I was in the network limbo area outside my house, where my phone struggles to transition from home WiFi to data. Worst of all: The failure would be horribly slow. I would have to drive fo…
It took me months to work out that "navigate to X" was the magic phrase to get google maps to do what you would expect in car navigation to do. The phrase that came more naturally to me was "give me directions to X", but that only gets you to the screen with the route and you still have to manually press the "start" button with your finger. And then it would randomly pick other modes of transport unless I remember to…
25 Years In Speech Technology and I still don’t talk to my computer
271–280 of 296 posts
Re: 25 Years In Speech Technology and I still don’t talk to my computer
#272I can imagine people using it exponentially more when you no longer get weird looks when you say “Ok google” or whatever in a supermarket.
Re: 25 Years In Speech Technology and I still don’t talk to my computer
#273Earlier quoted context omitted.
Discovery and browsing are not good on the assistant interface but I'd argue that's a constraint of voice vs visual. On a desktop/mobile, the screen holds the state so you can go back to previous entries, etc. Over voice, you mind holds the state which scales much much worse.
To the credit of the Shortcuts team at Apple, being able to visually select and define certain phrases for commonly completed tasks is helpful in this regard, but I’d guess only a couple % of the user base is even aware that Siri Shortcuts are possible.
Re: 25 Years In Speech Technology and I still don’t talk to my computer
#274Earlier quoted context omitted.
Rt. 5? Do you mean I-5? I'm curious if using "Rt." instead of "I-" is a regional thing. Where are you located? Typically Rt. would only be used for a small state road, not a major interstate highway.
It's used interchangeably for the number of any road, at least in my experience. It certainly gave appropriate responses when I tested it out.
Re: 25 Years In Speech Technology and I still don’t talk to my computer
#275Earlier quoted context omitted.
To be fair, that's just how Picard speaks (e.g. "engage"). I haven't noticed anyone else saying "Tea. Earl Grey. Hot". In any case I think this kind of speech is formulaic for the benefit of the audience, most of all, who are made aware through the formality that the speaker is addressing a machine. Additionally, we're watching navy men and women in space, so we expect them to speak to each other and to their compute…
Maybe I'm misremembering, but I distinctly remember that basically everyone ever shown interacting with a replicator follows the generic->specific parameter hierarchy.
Re: 25 Years In Speech Technology and I still don’t talk to my computer
#276Earlier quoted context omitted.
I live in Redmond, WA. Frankly, if someone asked to meet for a meal at eighteen, I would assume they would like to get together in a Microsoft cafeteria closest to the (what I believe to be non-existent) Building 18. My backup option would be to assume that they have received a hard blow to the head at some point in their life. Again, whether I would understand them or not, no one to my knowledge speaks like that in…
Officially building 18 does not exist. Anything you may have heard about building 18 is just rumour. If you have any questions regarding the purpose of building 18, you should direct them to Shelley in HR.
Re: 25 Years In Speech Technology and I still don’t talk to my computer
#277One thing I haven't seen discussed is the poor affordance / discoverability of speech technology. Google Home can do some clever things however, it also does not have the ability to do some very basic stuff. As a user, how do you know what Google Home can do and cannot do? It is just trial and error. And if Google Home introduces a new feature to be able to complete new types of queries. What now? How does a user kno…
Once they start hooking them up to conversational language models that can also submit queries, I think it is going to get a lot better. The conversational model results are starting to look very good.
Re: 25 Years In Speech Technology and I still don’t talk to my computer
#278Re: 25 Years In Speech Technology and I still don’t talk to my computer
#279Earlier quoted context omitted.
GUIs won over console workflows because GUIs have better discoverability and the "recall vs recognize" difference; it's mentally much easier to recognize the option you want when presented it than to recall the existence or the naming of that option. In those aspects of UX, voice interfaces have the same drawbacks as console apps when compared to a good GUI. Also, they have to work within the "bandwidth bottleneck" o…
This "recall vs recognize" point should be raised in every console vs GUI debate. It's pretty much the final word.
foo --
You're now present with a list of options and depending on your config, the man page one liner descriptionsRe: 25 Years In Speech Technology and I still don’t talk to my computer
#280Earlier quoted context omitted.
To the credit of the Shortcuts team at Apple, being able to visually select and define certain phrases for commonly completed tasks is helpful in this regard, but I’d guess only a couple % of the user base is even aware that Siri Shortcuts are possible.
Is this just normal phone Siri or HomePod Siri?