The term "human parity" refers to a comparison of the error rate, which is a single scalar summarizing performance in terms of mistakes made. It says nothing about the kind of mistakes, and I can easily imagine that machines qualitatively do not make at all the same kind of mistakes as humans. I'd be curious to know if the kind of mistakes machines make might strike human listeners as quite stupid, but maybe not.. ma…
The other day I asked Siri, "Remind me to pick up [daughter's name]" but it interpreted that as, "Remind me to pick up pussy." A more context-aware engine would realise we don't even have a cat. On the other hand it made me wonder to what extent Siri is trained on real-world data from other users.
Researchers reach human parity in conversational speech recognition
151–160 of 164 posts
Re: Researchers reach human parity in conversational speech recognition
#152Re: Researchers reach human parity in conversational speech recognition
#153The term "human parity" refers to a comparison of the error rate, which is a single scalar summarizing performance in terms of mistakes made. It says nothing about the kind of mistakes, and I can easily imagine that machines qualitatively do not make at all the same kind of mistakes as humans. I'd be curious to know if the kind of mistakes machines make might strike human listeners as quite stupid, but maybe not.. ma…
The term "human parity" refers to a comparison of the error The term 'human parity' refers to adding up all the bits in a human and taking the residue to a convenient modulus to detect human error.
Re: Researchers reach human parity in conversational speech recognition
#154Earlier quoted context omitted.
In my mind, this is analogous to the reasons why evaluation of lossy audio compression codecs requires human listening tests. Simply running some simplistic signal analysis like SNR (Signal-to-Noise Ratio) completely fails to capture the as-perceived quality of a compression implementation. To explore that analogy: In the case of lossy audio compression, the compressor deliberately introduces quantization noise into…
I've always evaluated audio codecs on differential signals - subtract A from A'. For telephony codecs, there are formal tests for MOS and/or PESQ .
These classic signal processing analyses tell you nothing about the correct operation of a psychoacoustic model codec design (i.e. MP3, AAC, Vorbis, etc.) An analysis of A (original signal) vs. A' (signal passed through a compression-decompression cycle) is just extracting the quantization noise introduced by the codec. That provides no information about how effective the codec was in masking that noise with the original signal content.
To illustrate, imagine a perversely designed codec: it runs two models, the first "good model" is a normal psychoacoustic model. The second "bad model" is the one the compressor uses: it applies the total amount of quantization noise allowed by the good model, but applies it in ways that are maximally annoying to human listeners. This isn't just avoiding masking, it's things like using noise correlated to the original signal, which is generally more obtrusive than uncorrelated noise, etc.
A codec using just the "good model" and one using the "good + bad model" would have (by definition) exactly the same introduced noise, but the latter would sound FAR worse to a human listener.
Re: Researchers reach human parity in conversational speech recognition
#155Earlier quoted context omitted.
The best case for animals demonstrating religious behavior (which is fundamentally different than believing in a god, mind you) is from a very convoluted interpretation of a food coma. http://bigthink.com/videos/can-animals-be-religious Outside of that considerable stretching, it doesn't exist. Also, I challenge your premise that conflates "belief" with "constructing narratives" Do you believe you will be alive tomor…
Do you alone in the world know what animals are thinking?
You provided... passive aggressiveness.
And if animals now engage in religious ritual, doesn't that make them stupid for not being atheist like you? I expect several YouTube videos of you trying to convert your cat into the enlightened ways of post-theism.
Re: Researchers reach human parity in conversational speech recognition
#156Earlier quoted context omitted.
A machine process qualifies as AI if it believes in a god. EDIT: I see the nuances of epistemological problems are lost to HN and knee-jerk culture war atheism still rules supreme.
FYI about your edit: You aren't being downvoted because of "knee-jerk culture war" or because people here are too dumb to understand your enlightened view of epistemology. Your original comment is an incredibly unique view with no explanation that people probably disagree with, and your edit makes you sound petulant, like you can't handle disagreement. That's the reason for the downvotes.
When you drag people out of their fantasies, they used to kill you. At least they downvote you these days, so I supposed that's better.
Yes, this is an incredibly unique view. Yes, people are going to react negatively to it because it fundamentally undermines their revenge fantasies. Yes, I can call them out on their unspoken biases no matter how uncomfortable that makes them.
Re: Researchers reach human parity in conversational speech recognition
#157Earlier quoted context omitted.
The best case for animals demonstrating religious behavior (which is fundamentally different than believing in a god, mind you) is from a very convoluted interpretation of a food coma. http://bigthink.com/videos/can-animals-be-religious Outside of that considerable stretching, it doesn't exist. Also, I challenge your premise that conflates "belief" with "constructing narratives" Do you believe you will be alive tomor…
>Also, I challenge your premise that conflates "belief" with "constructing narratives" I didn't conflate belief and narratives, I conflated belief with assigning high probability to those narratives. Sure, there are also simpler predictions we can estimate as well, but I think the thing that makes "belief in god" interesting from the perspective of judging intelligence is the complexity of the narrative. >Outside of…
And further more, if animals DO have religious behavior, then artificial intelligence research across the entire board is WOEFULLY inadequate to represent such a core part of the neurological interaction that generates religious behavior... a core behavioral capacity that was somehow missed in every single behaviorist research paper ever published since mankind mastered animal husbandry.
You either get in bed with the idea that only humans worship gods or you have to fundamentally throw all of psychology, sociology, and animal behavior studies completely out the window.
Re: Researchers reach human parity in conversational speech recognition
#158I'll admit I'm not very interested in speech recognition of this nature when it can't disambiguate the speaker. ie. the way Amazon Echo and other voice recognition systems can't tell the difference between a human in the room and the TV. Even when one might be clearly a female voice vs. a male. None of the voice recognition systems on the market learn my voice distinctly from my wife's or sons, and I don't want their…
for what it's worth, the "OK Google" functionality in my Android phone is trained against my voice and does a pretty good job at rejecting my wife's and coworkers commands.
My phone rarely listens to me until I hold down the home button but one guy I know, who has a slow, deep voice, triggers Google to start listening in normal conversation all the time.
Re: Researchers reach human parity in conversational speech recognition
#159Earlier quoted context omitted.
I've always evaluated audio codecs on differential signals - subtract A from A'. For telephony codecs, there are formal tests for MOS and/or PESQ .
Which, isn't substantially different than evaluating a codec on introduced SNR. These classic signal processing analyses tell you nothing about the correct operation of a psychoacoustic model codec design (i.e. MP3, AAC, Vorbis, etc.) An analysis of A (original signal) vs. A' (signal passed through a compression-decompression cycle) is just extracting the quantization noise introduced by the codec. That provides no i…
No joy.
I can't really follow your thinking above; sorry.
In audio, there is distortion, which is correlated with the signal. Noise is uncorrelated. Codec error would seem very much to me to be at least much more correlated than uncorrelated. MP3 artifacts sound at least more "phase-ey" than they sound anything like quantitization noise ( at least to me ). This may be because I have heard badly aligned tape machines make that sort of error happen. "Phase-ey" also triggers my (feeble) mind into thinking about allpass filters as the model for the error.
In terms of intelligibility, it's possible to improve intelligibility by adding noise, and by adding clipping - aviation comms does this at times. What destroys intelligibility is phenomes being destroyed by phase changes and bad amplitide errors ( where there are actually good amplitude errors ).
I've run an ABACUS voice quality analyzer several times, and I don't think it's purely a distortion analyzer - adding clipping at least can improve MOS/PESQ score surprisingly. Even more surprisingly, there's no mechanism available for calibrating gain staging on one.
Re: Researchers reach human parity in conversational speech recognition
#160Earlier quoted context omitted.
Which, isn't substantially different than evaluating a codec on introduced SNR. These classic signal processing analyses tell you nothing about the correct operation of a psychoacoustic model codec design (i.e. MP3, AAC, Vorbis, etc.) An analysis of A (original signal) vs. A' (signal passed through a compression-decompression cycle) is just extracting the quantization noise introduced by the codec. That provides no i…
I am not familiar with the term "introduced SNR"... Googling. No joy. I can't really follow your thinking above; sorry. In audio, there is distortion, which is correlated with the signal. Noise is uncorrelated. Codec error would seem very much to me to be at least much more correlated than uncorrelated. MP3 artifacts sound at least more "phase-ey" than they sound anything like quantitization noise ( at least to me ).…