Live data from Hacker News

25 Years In Speech Technology and I still don’t talk to my computer

matthewkaras.medium.com

181–190 of 296 posts

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#181
post #13

Earlier quoted context omitted.

> I can set my own alarms, thank you. And not only that, but a specialized user interface is often preferable to voice even when the assistant passes the Turing test. There's a reason that people use apps to get food delivery rather than calling.

> There's a reason that people use apps to get food delivery rather than calling. Well, if I had some type of personal assistant who worked for me (as in, a real human), I would just call out "Hey, Sam, please order pizza for me" and continue with whatever I was doing. The reason people don't make phone calls is they add a lot of additional friction. You have to dial the number, and wait for someone to pick up, and g…

Your assistant is probably smart enough to remember what kind of pizza you like and from where, and will just order that unless you tell them something else. In theory, there's no reason your phone auto-assistant couldn't do the same thing, but we seem to be a long way from any of them having that level of intelligence.

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#182

Earlier quoted context omitted.

Ouch. I just did a quickie test asking Siri to set me a reminder for 'tomorrow at eighteen'. No dice. It set it for 8 am. It doesn't even support 24 hour time - at least in English.

"...set an alarm for tomorrow @ eighteen hundred." works just fine. At least in U. S. English, I don't think I've ever heard anyone refer to 6 p. m. as "eighteen".

But that's the problem most people don't speak U.S. English as their primary language.

They use all kinds of languages with all kinds of accents and and ticks.

In my experience, and from the time I did a bit of NLP the situations is often along the line of. It works for mostly accent free simple English. Fails to get anywhere usable on most other languages and or accents. Sure that is to some degree because of missing training data. But for a consumer this doesn't change that for very many consumers this features work terrible bad.

Just out interest I tried out the youtube auto generated subtitle for a german video, but even at the parts where the text was super clear and well pronounced the result was hardly differentiable from randomly picking arbitrary words. It wasn't even that the algorithm choose similar sounding word. They where completely different words in many case. I think in a sentence of ~10 words in average 1 or 2 where correct. And that was at the parts where the text was unusually clear understandable. At other parts it wasn't even able to recognize that there where words...

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#183
post #70

Earlier quoted context omitted.

If you had two mics, you could probably work out a filter that captures audio roughly 'in front of the laptop' which would probably work well enough. But I think the wins are going to be in places where you don't normally have a computer, where a mouse and keyboard aren't natural companions to the task at hand. Yes, some environments will be noisy enough that speaking is a bad modality, but not all of them.

No, the point is that you'd be annoying the people around you if you were talking to your computer the whole time.

That is a lot of reading into the question that was actually asked: "how does it work?"

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#184
post #31

Earlier quoted context omitted.

Word recognition and sentence recognition. It's honestly quite shocking how sparse the research and implementations for everything is once you go beyond a single sentence/command that you shout at your personal assistants.

Frames used to be an idea in AI, but they seem to have been sidelined and possibly forgotten now. Frames mean that words and sentences have a context, and you can't understand conversations unless you understand the context. This starts from simple and obvious distinctions. E.g. - as a silly example - "make dinner" usually means "Prepare and cook an evening meal". But if you have a project called "dinner" it might me…

Oh, I didn't mean AGI levels of understanding, but even "simple" technical things that are likely building blocks necessary to get to that point like sentence boundary detection.

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#185
post #172

Earlier quoted context omitted.

"...set an alarm for tomorrow @ eighteen hundred." works just fine. At least in U. S. English, I don't think I've ever heard anyone refer to 6 p. m. as "eighteen".

but you would still understand them if they did or take 5 seconds to ask a clarifying question.

I live in Redmond, WA. Frankly, if someone asked to meet for a meal at eighteen, I would assume they would like to get together in a Microsoft cafeteria closest to the (what I believe to be non-existent) Building 18. My backup option would be to assume that they have received a hard blow to the head at some point in their life.

Again, whether I would understand them or not, no one to my knowledge speaks like that in U. S. English. It is a great example to use to show the quirks of language. It is a bad example to use to show that Siri "doesn't even support 24 hour time - at least in English".

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#186
post #75

Earlier quoted context omitted.

It might be a good way to learn to talk to people, or just talk... Actually, just use Discord for that, join random servers and the voice chats.

That sounds even worse.

No no, it actually works. I am working from home most of the time, and chatting on Discord with random strangers helps me keep my speaking skills, and just stay sane(ish).

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#187
post #48

Earlier quoted context omitted.

Well, to extend the GUI/console metaphor, it means that at some point soon, we'll all be using NLP because it's dramatically more user-friendly for the vast majority of people.

I understand your point, but I'm not sure that GUIs won out because they were dramatically more user-friendly. It certainly helped, but I think they won because it made multitasking possible. Multitasking from the users perspective that is, the ability to interact with more than one application at the same time. That was just not possible on a console, so even people who didn't need user friendliness were able to do…

> That was just not possible on a console

You may be thinking of DOS, which yes had almost no multitasking ability available.

However there were multiple timesharing operating systems that existed before the PC and GUIs, Unix being the most famous and still around.

Multitasking is quite possible on a Linux console for example. It has 5 or more consoles, each handling different users, each being able to be split via screen/tmux. Each shell can run jobs in the background as well.

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#188
The main problem now is not speech recognition. It's a kind of uncanny valley effect.

You can speak to these assistants, but the language is still restricted. They show little to no common sense. It's a lot of party tricks bundled together.

You can't interrupt them and it's hard to correct them.

On the speech recognition side, an issue I've found (although is a rather niche one) is triggered because I'm bilingual (I'm fluent in spanish and english).

Speech recognition only works well on a single language.

I have Alexa set to speak english, for example when I'm searching for a song with a spanish title, I have to try to fudge the name into a fake english pronunciation for it to produce the right phonemes that will match the song title, rather than just say the name properly.

Also, if it misses the match, there's no easy way to stop it and say "No, not that one", and be presented with a list of similar matches.

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#189
post #5

A big problem with the assistants is that as soon as they fail at a query they seem stupid, I feel stupid, and I stop using them for a long time. Last weekend I had the following failed queries: "OK Google, what is the air quality like at Mt. Shasta today?" "OK Google, add a waypoint for the last gas station before the mountain pass" "OK Google, what percentage of people can you detect to be wearing masks on recent I…

If google assistant or whatever worked like a person of average intelligence that I could talk to and get to do stuff for me, that would be incredible. I've been on a roadtrip and your examples made me realize how much of a time/attention saver something like that would have been.

Re: 25 Years In Speech Technology and I still don’t talk to my computer

#190
post #188

The main problem now is not speech recognition. It's a kind of uncanny valley effect. You can speak to these assistants, but the language is still restricted. They show little to no common sense. It's a lot of party tricks bundled together. You can't interrupt them and it's hard to correct them. On the speech recognition side, an issue I've found (although is a rather niche one) is triggered because I'm bilingual (I'…

I don't think it's a niche issue outside of the United States, in countries where English is widely known. It makes it pretty much impossible to use Siri on Apple TV for example, because so many movies or TV series are named in English but also in the local language.
Post reply on HN