Live data from Hacker News

25 Years In Speech Technology and I still don’t talk to my computer

matthewkaras.medium.com

21–30 of 296 posts

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#22

Very unlikely I will ever talk to my computer irrespective of how good the speed technology gets. If I have to talk to my computer how will I work in crowded places ? How does it work ?

Even at home without any risk of disturbing coworkers, it would have to be extremely intuitive to offer any speed advantage.

I can open Microsoft Word faster than I can say "open microsoft word". It would have to be smart enough to short circuit the entire process of doing something useful with Word.

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#24
In my opinion, the speech-driven computing in Star Trek is the best depiction of the technology.

The interaction is very fluid, low-latency, and accurate, and the system doesn't force itself on you: there are still plenty of non-speech-based user interfaces to be found all over a starship.

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#25
Question: Can I write my own assistant on mobile devices that is an actual assistant? As in, can listen for the key activation word in the background? Like the "official" assistants.

As far as I can tell, the assistant APIs seem to be like plugins? On Android for example, it appears custom assistants still run through Google assistant.

I want to be able to say "TriggerWord, do X and Y" and the OS activates my app, passes the voice sample and I take care of all the language processing from there. Which doesn't seem possible...

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#26
post #5

A big problem with the assistants is that as soon as they fail at a query they seem stupid, I feel stupid, and I stop using them for a long time. Last weekend I had the following failed queries: "OK Google, what is the air quality like at Mt. Shasta today?" "OK Google, add a waypoint for the last gas station before the mountain pass" "OK Google, what percentage of people can you detect to be wearing masks on recent I…

When they do work though you feel like you’re at the helm of the future.

I’ve figured out what queries almost never seem to fail and use those almost exclusively. I don’t get creative.

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#27
This is a great example of "it's not my preferred mode, so it must not be anyone else's."

Dictation is widely used in medical transcription.

Dictation is a killer way to write a first draft quickly, transcribe rough written notes after a meeting, etc. Also, about half of my emails are dictated, and I know I'm not the only one. It takes some time to get used to, but once you're there (like touch typing!), you can't go back.

..etc..

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#28
post #5

A big problem with the assistants is that as soon as they fail at a query they seem stupid, I feel stupid, and I stop using them for a long time. Last weekend I had the following failed queries: "OK Google, what is the air quality like at Mt. Shasta today?" "OK Google, add a waypoint for the last gas station before the mountain pass" "OK Google, what percentage of people can you detect to be wearing masks on recent I…

Well put.

I don't know how much of FAANG's budget goes towards improving voice assistants, but considering how much cash these companies have on hand and their operating budgets (the size of some smaller European countries), the progress in that area is just super disappointing.

The 3 most common use cases were refined a long time ago (Directions, Alarms and Play Music) and everything else ran into a hard wall.

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#29
post #14
post #5

A big problem with the assistants is that as soon as they fail at a query they seem stupid, I feel stupid, and I stop using them for a long time. Last weekend I had the following failed queries: "OK Google, what is the air quality like at Mt. Shasta today?" "OK Google, add a waypoint for the last gas station before the mountain pass" "OK Google, what percentage of people can you detect to be wearing masks on recent I…

“Alexa, play punk rock” “Playing ” “Alexa stop. Alexa play punk rock playlist” “Cannot find punk rock playlist” “Alexa play punk rock 00’s playlist” “Can’t find” “Alexa play early 2000’s punk rock” “Playing punk rock 00’s playlist” It’s like she’s trying to mock me.

Burglar: "Put your hands up and show me where the money is, I won't hurt you..."

Me: "Alright man, I'll tell you where is money ... ALEXA CALL THE POLICE!"

Alexa: "Shuffling songs by The Police"

* EVERY BREATH YOU TAKE plays as I get punched 24 times *

from https://twitter.com/ppathole/status/1092034892249079813

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#30
post #5

A big problem with the assistants is that as soon as they fail at a query they seem stupid, I feel stupid, and I stop using them for a long time. Last weekend I had the following failed queries: "OK Google, what is the air quality like at Mt. Shasta today?" "OK Google, add a waypoint for the last gas station before the mountain pass" "OK Google, what percentage of people can you detect to be wearing masks on recent I…

But your examples are way harder than they sound. Speech or non speech analysers have a hard time with context. What do you mean by "recent" photos. And what percentage? Of people wearing mask in each photo, or of 1 or more people with a mask in the whole set of photos? Or the percentage of photo having all people wearing a mask. We humans make a lot of deduction from context. We haven't been able to teach computers this aspect for 25 years. It has started more recently, deep learning shows potential.
Post reply on HN