Live data from Hacker News

25 Years In Speech Technology and I still don’t talk to my computer

matthewkaras.medium.com

121–130 of 296 posts

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#121
post #5

A big problem with the assistants is that as soon as they fail at a query they seem stupid, I feel stupid, and I stop using them for a long time. Last weekend I had the following failed queries: "OK Google, what is the air quality like at Mt. Shasta today?" "OK Google, add a waypoint for the last gas station before the mountain pass" "OK Google, what percentage of people can you detect to be wearing masks on recent I…

Actually, ability to delete all iPhone alarms at once via Siri is a life saver. I know no other way to bulk delete l/disable alarms in iOS.

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#122
post #13

Earlier quoted context omitted.

> I can set my own alarms, thank you. And not only that, but a specialized user interface is often preferable to voice even when the assistant passes the Turing test. There's a reason that people use apps to get food delivery rather than calling.

It's funny, but for years I've felt like "Hey Siri, set an alarm for 7am" was 1000x easier than using the clunky Clock UI and it's almost exclusively what I use Siri for. Tasks that are so perfectly well-defined are exactly what this "smart" tech is useful for. Except that recently, Siri screwed me. I said "hey Siri, set an alarm for 6:30am" and her response was "Ok, your 6:30am alarm is on," but what she actually me…

Alarms and reminders are my most frequent uses of Siri (and to entertain/frustrate the kids...). It's also pretty good for hands-free quick texts while driving and hands-free calling, though I don't do those often.

Your alarm example is frustrating. Of course, you could say "set alarm for tomorrow at 6:30 am" which would work, but then you're back in the realm not of natural language commands, but formal commands that exist in the uncanny valley, just similar enough to natural language to be irritating when they fail.

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#123
post #13
post #5

A big problem with the assistants is that as soon as they fail at a query they seem stupid, I feel stupid, and I stop using them for a long time. Last weekend I had the following failed queries: "OK Google, what is the air quality like at Mt. Shasta today?" "OK Google, add a waypoint for the last gas station before the mountain pass" "OK Google, what percentage of people can you detect to be wearing masks on recent I…

> I can set my own alarms, thank you. And not only that, but a specialized user interface is often preferable to voice even when the assistant passes the Turing test. There's a reason that people use apps to get food delivery rather than calling.

The only two questions I ever ask Siri are "how cold is it outside?" and "wake me up {at HHam, in X minutes}"

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#124
post #5

A big problem with the assistants is that as soon as they fail at a query they seem stupid, I feel stupid, and I stop using them for a long time. Last weekend I had the following failed queries: "OK Google, what is the air quality like at Mt. Shasta today?" "OK Google, add a waypoint for the last gas station before the mountain pass" "OK Google, what percentage of people can you detect to be wearing masks on recent I…

But your examples are way harder than they sound. Speech or non speech analysers have a hard time with context. What do you mean by "recent" photos. And what percentage? Of people wearing mask in each photo, or of 1 or more people with a mask in the whole set of photos? Or the percentage of photo having all people wearing a mask. We humans make a lot of deduction from context. We haven't been able to teach computers…

Exactly, which is why these assistants are not very useful beyond simple tasks you could just do yourself. If they're only good at things that are easy for you to do, then what's the point of them besides not needing your hands to do simple tasks?

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#125
post #13

Earlier quoted context omitted.

> I can set my own alarms, thank you. And not only that, but a specialized user interface is often preferable to voice even when the assistant passes the Turing test. There's a reason that people use apps to get food delivery rather than calling.

> There's a reason that people use apps to get food delivery rather than calling. Well, if I had some type of personal assistant who worked for me (as in, a real human), I would just call out "Hey, Sam, please order pizza for me" and continue with whatever I was doing. The reason people don't make phone calls is they add a lot of additional friction. You have to dial the number, and wait for someone to pick up, and g…

Pizza is a weird example because pizza joint menus are more-or-less standard and unsurprising. For any non-pizza restaurant I'll want to start by scanning a menu, which is way more efficient than listening to a menu be read to me.

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#126

It's also the unprompted responses that has started to bug me. I don't know if Google is having a bad rollout or this is deliberate but my Google Home is being triggered a lot more often now and I don't recall anything remotely close to the wake phrase being said. Also, the responses to questions when actually prompted are not at all what I'm expecting. Particularly, Google has trouble understanding a lot of context…

I had this problem a lot with Cortana popping up during meetings and seizing control of my microphone. In the end I spent probably 15 minutes trying to figure out how to turn off voice activation because Microsoft doesn't make it easy.

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#127
post #38

Earlier quoted context omitted.

Even with directions, it fails miserably IMO. The most it seems to be able to do is "navigate to X". I want: - "Take me on the most scenic route to X." Can't you figure that out from social media tags? Simple first order solution: routes that have more photos with more likes = more scenic. Took me 1 minute to think of that. And 1000 engineers at Google couldn't implement that? These data crunching tasks are the kind…

Scenic: Garmin's devices appear to do some calculation on number of times the road doesn't go straight over a given distance. Seems to work well enough. OTOH, someone at Google has to make this a feature. Paved/unpaved: that's one suck-ass assistant you've got there. I know the Garmin on the dash of my motorcycle will give that option. The Garmin RV-specific GPS has loads of other options, such as avoid any low overp…

Yeah, I would have expected Google to even take it to the next level: Ask what model of car the user has (better yet! Identify it from a combination of the car's Bluetooth MAC address and a machine learning model trained on the audio spectrum of the engine noise). Look at where all cars of different types are able to travel, where they make U-turns and turn around, and at what GPS locations they make calls to roadside assistance numbers. Incorporate weather data as well, e.g. what season, and whether it has rained lately. Between all of that data you should be able to advise whether a given vehicle should be able to traverse a given road on a given day of the year.

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#128

I talk to my phone in the car all the time. Hands free. "send message to john" "how far is it to" "get directions to" "play podcast" "play audiobook"

And then your phone transfers the recorded sounds to someone else's computer to do the work. You still don't talk to your computer phone. You talk through it.

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#129
post #71

One thing I haven't seen discussed is the poor affordance / discoverability of speech technology. Google Home can do some clever things however, it also does not have the ability to do some very basic stuff. As a user, how do you know what Google Home can do and cannot do? It is just trial and error. And if Google Home introduces a new feature to be able to complete new types of queries. What now? How does a user kno…

Agreed. More broadly, I'd say that no one has made a good UI yet. The mac/lisa/star had a UI that people could learn. iOS...

In some ways a voice UI has bigger problems to deal with than PC GUIs or iOS. Those UIs were replacing pre-existing UIs (eg blackberry, dos, unix, norton) and they could target whatever tasks a smartphone/PC needed to do. For voice UIs, it's a cold start. It's not even obvious what an audio only computer should do. Our mental model for a "virtual assistant" is a person-2-person exchange, and computers still aren't great at communicating like people.

FWIW, I think slipping into existing niches is the way to go. That's where a useful voice ui will be discovered. Car stuff, accessibility software, living room controls.. At least these have clear goals. Voice operating spotify, netflix or just an iphone is something people actually need and will use if its useful.

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#130
post #69

Earlier quoted context omitted.

I think you underestimate the resistance. The problem isn't just that people don't want to do it, it's also that people don't want other people to do it because it's annoying. It might pass in a coffee shop, it probably won't in an airplane, it will almost certainly never pass in a library. People will try it regardless of the appropriateness of the location, because some people aren't aware or are assholes, and as a…

People talk to each other all the time in coffee shops. If talking to your computer becomes as natural as talking to a friend, why won't it be acceptable in a shop?

People talk to each other all the time in airplanes too, and even libraries, but it turns out that people are both more understanding of people talking to other humans currently present (rather than say on a cell phone) and people are better at talking to other humans currently present respectfully than they are at talking to devices (or at least people via devices, but I think we can extrapolate).

I agree if it was already the social norm it wouldn't be a problem (same with Google glass), but it turns out that the technology being ready isn't always enough to make the social norms change.

Post reply on HN