Live data from Hacker News

25 Years In Speech Technology and I still don’t talk to my computer

matthewkaras.medium.com

61–70 of 296 posts

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#61
post #38

Earlier quoted context omitted.

Well put. I don't know how much of FAANG's budget goes towards improving voice assistants, but considering how much cash these companies have on hand and their operating budgets (the size of some smaller European countries), the progress in that area is just super disappointing. The 3 most common use cases were refined a long time ago (Directions, Alarms and Play Music) and everything else ran into a hard wall.

Even with directions, it fails miserably IMO. The most it seems to be able to do is "navigate to X". I want: - "Take me on the most scenic route to X." Can't you figure that out from social media tags? Simple first order solution: routes that have more photos with more likes = more scenic. Took me 1 minute to think of that. And 1000 engineers at Google couldn't implement that? These data crunching tasks are the kind…

Scenic: Garmin's devices appear to do some calculation on number of times the road doesn't go straight over a given distance. Seems to work well enough. OTOH, someone at Google has to make this a feature.

Paved/unpaved: that's one suck-ass assistant you've got there. I know the Garmin on the dash of my motorcycle will give that option. The Garmin RV-specific GPS has loads of other options, such as avoid any low overpasses that will rip the solar panels off my RV. Though it doesn't look like Apple is any better than Google, only because Apple Maps doesn't offer the option. Anyway, satellite and street view? Rand McNally had this information since dirt was first created, no need to get the satellites out.

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#62
post #31
post #23

We've been pretty good at word recognition for a while - speech not so much. Conflating the two has lead to a lot of confusion.

Word recognition and sentence recognition. It's honestly quite shocking how sparse the research and implementations for everything is once you go beyond a single sentence/command that you shout at your personal assistants.

That's fair, NLP has done decently with mechanical deconstruction of normal sentences for quite a while now. But as you note, mapping that onto a template for response is a long way from "understanding".

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#63
post #31
post #23

We've been pretty good at word recognition for a while - speech not so much. Conflating the two has lead to a lot of confusion.

Word recognition and sentence recognition. It's honestly quite shocking how sparse the research and implementations for everything is once you go beyond a single sentence/command that you shout at your personal assistants.

Frames used to be an idea in AI, but they seem to have been sidelined and possibly forgotten now.

Frames mean that words and sentences have a context, and you can't understand conversations unless you understand the context.

This starts from simple and obvious distinctions. E.g. - as a silly example - "make dinner" usually means "Prepare and cook an evening meal". But if you have a project called "dinner" it might mean "build and compile 'dinner'" An AGI should be able to understand the difference, and ask for clarification if it doesn't.

Eventually you end up with subtextual and implied communication - e.g. "I'm fine" can mean two completely opposite things depending on tone of voice and the contents of minutes-to-years of previous conversations.

All of this is many orders of magnitude harder to handle than "Bedroom lights off."

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#64
post #47

To avoid wrist pain, I have started dictating quite a lot as a way to enter large amounts of text into my laptop and phone. It's really very good, especially on android. Aside from the relatively large downside of it being noisy, I like it almost as much as I did typing.

Do you find yourself going back and revising what you've dictated? When I type, I'm frequently pausing, going to other parts of the document, and deleting things I've already written. All of these actions I find more annoying when dictating. Not to mention some of the baffling Random capitalization Choices and, over/under insertion of punctuation.

In general, I find dictation useful on my phone to take notes to myself. And maybe to send a message in a chat situation. But it just doesn't work for me for anything slightly more formal / long.

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#65
post #5

A big problem with the assistants is that as soon as they fail at a query they seem stupid, I feel stupid, and I stop using them for a long time. Last weekend I had the following failed queries: "OK Google, what is the air quality like at Mt. Shasta today?" "OK Google, add a waypoint for the last gas station before the mountain pass" "OK Google, what percentage of people can you detect to be wearing masks on recent I…

But would you have gotten better results if you had typed those same queries into some website?

While an artificial general intelligence would be capable of fluent speech, a perfectly good speech system does not necessarily need to be an AGI.

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#66
post #5

A big problem with the assistants is that as soon as they fail at a query they seem stupid, I feel stupid, and I stop using them for a long time. Last weekend I had the following failed queries: "OK Google, what is the air quality like at Mt. Shasta today?" "OK Google, add a waypoint for the last gas station before the mountain pass" "OK Google, what percentage of people can you detect to be wearing masks on recent I…

But your examples are way harder than they sound. Speech or non speech analysers have a hard time with context. What do you mean by "recent" photos. And what percentage? Of people wearing mask in each photo, or of 1 or more people with a mask in the whole set of photos? Or the percentage of photo having all people wearing a mask. We humans make a lot of deduction from context. We haven't been able to teach computers…

Context is hard, but it seems like “recent” means (99% of the time) order by date descending, grab the first 15 or so, and then how many of those photos contain a person with a mask.

Maybe the difficult part is whether or not you look for the 15 most recent photos containing people, or the 15 most recent photos of anything.

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#67

Question: Can I write my own assistant on mobile devices that is an actual assistant? As in, can listen for the key activation word in the background? Like the "official" assistants. As far as I can tell, the assistant APIs seem to be like plugins? On Android for example, it appears custom assistants still run through Google assistant. I want to be able to say "TriggerWord, do X and Y" and the OS activates my app, pa…

https://snowboy.kitt.ai/ is about the only good third party hotword detection engine. You can hook it into Assistant with a bit of work, and with Tasker/Automate and a rooted phone, you can open apps, press buttons, pass voice commands, and more.

Imo, Assistant is limited because of privacy, see how it asks you to opt in for a more personalized experience and it still won't unlock your phone for you.

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#68

I've had a similar experience, worked in voice for a long time. I simply have no desire whatsoever to speak to my electronics, I don't find it helpful or useful in anyway.

I'm not quite that extreme, but I'd use voice if it was at least as smart as a human. And it isn't - yet. "Turn on the lights" isn't exciting. "Make dinner and do the laundry" would be exciting, but that's at least 25 years and some major advances in robotics away. IoT is very crude and contrived compared to what would be possible with an active technology that could do useful physical things of all kinds.

For something like that I'd rather press a button or have it happen on a timer.

Where it could really shine is in more rare commands. "Make me some Pad Thai". Unless you're a huge fan, you wont want a button for this. And it's faster to say than type into your phone.

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#69
post #19
post #6

Earlier quoted context omitted.

I think the real problem is social, no one wants to hear you talking at your computer in a coffee shop.

That would quickly change if people actually wanted to do it though. Social norms are just unwritten agreements - if enough people decide to dictate their novels and blog posts in coffeeshops then the default will quickly change so talking is acceptable.

I think you underestimate the resistance. The problem isn't just that people don't want to do it, it's also that people don't want other people to do it because it's annoying.

It might pass in a coffee shop, it probably won't in an airplane, it will almost certainly never pass in a library. People will try it regardless of the appropriateness of the location, because some people aren't aware or are assholes, and as a result the entire technology will get a bad rap. See google glass as an example.

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#70

Very unlikely I will ever talk to my computer irrespective of how good the speed technology gets. If I have to talk to my computer how will I work in crowded places ? How does it work ?

If you had two mics, you could probably work out a filter that captures audio roughly 'in front of the laptop' which would probably work well enough. But I think the wins are going to be in places where you don't normally have a computer, where a mouse and keyboard aren't natural companions to the task at hand. Yes, some environments will be noisy enough that speaking is a bad modality, but not all of them.

No, the point is that you'd be annoying the people around you if you were talking to your computer the whole time.
Post reply on HN