Live data from Hacker News

Researchers reach human parity in conversational speech recognition

blogs.microsoft.com

51–60 of 164 posts

Re: Researchers reach human parity in conversational speech recognition

#51

The term "human parity" refers to a comparison of the error rate, which is a single scalar summarizing performance in terms of mistakes made. It says nothing about the kind of mistakes, and I can easily imagine that machines qualitatively do not make at all the same kind of mistakes as humans. I'd be curious to know if the kind of mistakes machines make might strike human listeners as quite stupid, but maybe not.. ma…

That reminds me of the photocopier that used a very poorly chosen image compression algorithm. It would sometimes substitute similar looking characters to save space. A $7000 figure in a budget might become $9000 and a person looking over the document would have no indication that the '9' had been cut and pasted from somewhere else on the page. Google Voice transcription often makes easily detectable mistakes that I…

For anyone else wondering about that copier: http://www.theregister.co.uk/2013/08/06/xerox_copier_flaw_me...

Re: Researchers reach human parity in conversational speech recognition

#52
post #40
post #28

I look forward to being able to converse with Microsoft's research team as easily as I can with humans. I hope that one day, journalists can learn to write headlines with similarly low rates of error.

It's on purpose... journalists learned to do it this way.

http://www.smbc-comics.com/comics/20090830.gif

Re: Researchers reach human parity in conversational speech recognition

#53
post #44

Earlier quoted context omitted.

Being made to believe in a god isn't the same as believing in one. Silicon Valley thinks AI is not an epistemological problem, as if neurons can be perfectly simulated atomically and that all intelligence processes can be categorized as structured vs. unstructured. Very naive conclusions. Ironically, they BELIEVE if you simulate the axiomatic neuron perfectly, emergent properties of intelligence will mystically emerg…

For Artificial Intelligence to believe in God would be for it to believe in its makers, which are human. A machine would be intelligent to be aware that it was made by man. A man doesn't need to believe in God to be intelligent. In fact, the two are pretty much inversely related.

>For Artificial Intelligence to believe in God would be for it to believe in its makers //

That's a very limited sense of the idea of a god, and certainly doesn't match with definitions of God [a singular, eternal, omnipotent, omniscient, deity] that I've come across.

Merely making something doesn't make you a god, not even if that thing appears to display intelligence. Some sense of one of the characteristics of existing in a separate spiritual realm, having power/knowledge beyond that possible in the present realm, having an existence that's not bounded (eg physically) within the normally experienced space of the "mortals". They seem like a start for basic level definitions of a god.

Re: Researchers reach human parity in conversational speech recognition

#54
post #12

Earlier quoted context omitted.

I don't get why that's frustrating, without image and speech recognition AI isn't possible. -- How intelligent could people be without sensory perception? Hellen Keller is the exception, take away the ability for humans to process sound, and images as a species and we wouldn't be nearly as advanced as we are today even if we still had the same brain structure and mental capacity. Ray Kurzweil understood this and is w…

Perhaps because "true AI" is an illusion. There is always a better term that has meaning. And yet people like Ray Kurzweil toss the term around as if it does have meaning. What's the point unless you define it? You might as well reference anti gravity or teleportation.

It's something to work towards, always on the horizon.

Kurzweil was a great inventor once but he's a prophet now and he needs to speak prophetically, not scientifically.

Re: Researchers reach human parity in conversational speech recognition

#55

Earlier quoted context omitted.

I can think of no better test of whether some specific being is intelligent than fooling another human into thinking they are. Can you think of a better one? > because it's a behavioral or functional test What else would you test for in an AI other than its output and behavior? Conversely: how would you test a human for intelligence other than through its output and behavior?

Surely that depends upon the human? An infant would do a terrible job. An institutionalized vegetable would fail to provide meaningful information. A not-very-bright cousin of mine has been heard responding to robo-calls. Would his opinion do? I think 'intelligence' is ambiguous, but many parts of it can be measured to some degree. Lets give the AI the SAT test perhaps? Or a test of hypothetical arguments. Or ask it…

A test that confuses a human for a machine is not a problem. False negatives don't remove value from this test.

Re: Researchers reach human parity in conversational speech recognition

#56
post #8

Earlier quoted context omitted.

The touring test shows AI has little to do with image or speech recognition. An AI that operates in a purely virtual environment would still be revolutionary.

The Turing test doesn't test for intelligence , simply human-like communication. Things like lying, typing errors, or reaction to insult aren't indicators of an artificial intelligence, but are requirements for passing the Turing test. One of the biggest criticisms of the Turing test is, because it's a behavioral or functional test, there is no way to ensure that the computer is actually thinking intelligently at all…

It directly tests for intelligence, you confuse false negatives with false positives. An intelligent person that fails the test because they don't know a common language is fine.

Note: Medical tests generate both false positives and false negatives, they are still useful.

Re: Researchers reach human parity in conversational speech recognition

#58
post #3

The actual paper has a section on error analysis that is particularly enlightening: https://arxiv.org/abs/1610.05256 On the CallHome dataset humans confuse words 4.1% of the time, but delete 6.5% of words, most commonly deleting the word "I". Their ASR system confuses 6.5% of words on this dataset, but only deletes 3.3% of words, so depending on how you view this their claim about being better than humans isn't defin…

I'm not an expert on this at all, but my suspicion is that humans are more tolerant of missing words than we are of mistaken words. Especially a word like "I", we just assume it if it's missing.

So if this speech recognition is for creating a transcript for humans, this (in my uneducated opinion) isn't as good as humans do, at least not yet.

Re: Researchers reach human parity in conversational speech recognition

#59
post #3

The actual paper has a section on error analysis that is particularly enlightening: https://arxiv.org/abs/1610.05256 On the CallHome dataset humans confuse words 4.1% of the time, but delete 6.5% of words, most commonly deleting the word "I". Their ASR system confuses 6.5% of words on this dataset, but only deletes 3.3% of words, so depending on how you view this their claim about being better than humans isn't defin…

I was recently at a bar where they showed a movie with incomprehensible subtitles (English to English). I assume this was because they skimped and bought automatic subtitling.

I think one important aspect is while humans miss words, they often get the sentence meaning correct. When computers miss words, they tend to substitute words that sound similar. That's readable if you have time but not necessarily as a stream of text going by...

Re: Researchers reach human parity in conversational speech recognition

#60

The term "human parity" refers to a comparison of the error rate, which is a single scalar summarizing performance in terms of mistakes made. It says nothing about the kind of mistakes, and I can easily imagine that machines qualitatively do not make at all the same kind of mistakes as humans. I'd be curious to know if the kind of mistakes machines make might strike human listeners as quite stupid, but maybe not.. ma…

Section 9 in the paper[1] is all about comparing these mistakes between the system and humans. The most common mistakes for humans and the system are in tables 9—11.

We find that the artificial errors are substantially the same as human ones with one large exception confusions between backchannel words [acknowledgment words like “uh-huh”] and hesitations.

The difference they found, but suspect might be a result of the different transcription guidelines of the training corpus: we see that by far the most common error in the ASR system is the confusion of a hesitation in the reference for a backchannel in the hypothesis. People do not seem to have this problem.

[1]: https://arxiv.org/abs/1610.05256

Post reply on HN