Live data from Hacker News

Researchers reach human parity in conversational speech recognition

blogs.microsoft.com

151–160 of 164 posts

Re: Researchers reach human parity in conversational speech recognition

#151

The term "human parity" refers to a comparison of the error rate, which is a single scalar summarizing performance in terms of mistakes made. It says nothing about the kind of mistakes, and I can easily imagine that machines qualitatively do not make at all the same kind of mistakes as humans. I'd be curious to know if the kind of mistakes machines make might strike human listeners as quite stupid, but maybe not.. ma…

The other day I asked Siri, "Remind me to pick up [daughter's name]" but it interpreted that as, "Remind me to pick up pussy." A more context-aware engine would realise we don't even have a cat. On the other hand it made me wonder to what extent Siri is trained on real-world data from other users.

When SIRI reminds you to go grab that p-ssy, it's finally human enough to run for president.

Re: Researchers reach human parity in conversational speech recognition

#153

The term "human parity" refers to a comparison of the error rate, which is a single scalar summarizing performance in terms of mistakes made. It says nothing about the kind of mistakes, and I can easily imagine that machines qualitatively do not make at all the same kind of mistakes as humans. I'd be curious to know if the kind of mistakes machines make might strike human listeners as quite stupid, but maybe not.. ma…

The term "human parity" refers to a comparison of the error The term 'human parity' refers to adding up all the bits in a human and taking the residue to a convenient modulus to detect human error.

Sounds messy ;)

Re: Researchers reach human parity in conversational speech recognition

#154

Earlier quoted context omitted.

In my mind, this is analogous to the reasons why evaluation of lossy audio compression codecs requires human listening tests. Simply running some simplistic signal analysis like SNR (Signal-to-Noise Ratio) completely fails to capture the as-perceived quality of a compression implementation. To explore that analogy: In the case of lossy audio compression, the compressor deliberately introduces quantization noise into…

I've always evaluated audio codecs on differential signals - subtract A from A'. For telephony codecs, there are formal tests for MOS and/or PESQ .

Which, isn't substantially different than evaluating a codec on introduced SNR.

These classic signal processing analyses tell you nothing about the correct operation of a psychoacoustic model codec design (i.e. MP3, AAC, Vorbis, etc.) An analysis of A (original signal) vs. A' (signal passed through a compression-decompression cycle) is just extracting the quantization noise introduced by the codec. That provides no information about how effective the codec was in masking that noise with the original signal content.

To illustrate, imagine a perversely designed codec: it runs two models, the first "good model" is a normal psychoacoustic model. The second "bad model" is the one the compressor uses: it applies the total amount of quantization noise allowed by the good model, but applies it in ways that are maximally annoying to human listeners. This isn't just avoiding masking, it's things like using noise correlated to the original signal, which is generally more obtrusive than uncorrelated noise, etc.

A codec using just the "good model" and one using the "good + bad model" would have (by definition) exactly the same introduced noise, but the latter would sound FAR worse to a human listener.

Re: Researchers reach human parity in conversational speech recognition

#155

Earlier quoted context omitted.

The best case for animals demonstrating religious behavior (which is fundamentally different than believing in a god, mind you) is from a very convoluted interpretation of a food coma. http://bigthink.com/videos/can-animals-be-religious Outside of that considerable stretching, it doesn't exist. Also, I challenge your premise that conflates "belief" with "constructing narratives" Do you believe you will be alive tomor…

Do you alone in the world know what animals are thinking?

I provided the very best and most recent research for the absurd position that animals supposedly have their own animal gods, attend regular communal worshipping ceremonies, engage in animal religious ritual, and demonstrate religious behavior outside of a priori instinct and social fitnesses.

You provided... passive aggressiveness.

And if animals now engage in religious ritual, doesn't that make them stupid for not being atheist like you? I expect several YouTube videos of you trying to convert your cat into the enlightened ways of post-theism.

Re: Researchers reach human parity in conversational speech recognition

#156
post #121

Earlier quoted context omitted.

A machine process qualifies as AI if it believes in a god. EDIT: I see the nuances of epistemological problems are lost to HN and knee-jerk culture war atheism still rules supreme.

FYI about your edit: You aren't being downvoted because of "knee-jerk culture war" or because people here are too dumb to understand your enlightened view of epistemology. Your original comment is an incredibly unique view with no explanation that people probably disagree with, and your edit makes you sound petulant, like you can't handle disagreement. That's the reason for the downvotes.

I'm being downvoted because my comment spat in the face of those who think general artificial intelligence automatically implies some superior Nietzschean post-human that is qualified to run the perfect communist utopia since it no longer has distracting human emotions like greed or envy. Somewhere, a Sanfraninite spurted himself to completion while watching a live feed of indigenous humans tending to some noble savage rice field he rented on AgroBnb with Bitcoin, mined entirely by solar powered GPUs made out of free trade copper.

When you drag people out of their fantasies, they used to kill you. At least they downvote you these days, so I supposed that's better.

Yes, this is an incredibly unique view. Yes, people are going to react negatively to it because it fundamentally undermines their revenge fantasies. Yes, I can call them out on their unspoken biases no matter how uncomfortable that makes them.

Re: Researchers reach human parity in conversational speech recognition

#157

Earlier quoted context omitted.

The best case for animals demonstrating religious behavior (which is fundamentally different than believing in a god, mind you) is from a very convoluted interpretation of a food coma. http://bigthink.com/videos/can-animals-be-religious Outside of that considerable stretching, it doesn't exist. Also, I challenge your premise that conflates "belief" with "constructing narratives" Do you believe you will be alive tomor…

>Also, I challenge your premise that conflates "belief" with "constructing narratives" I didn't conflate belief and narratives, I conflated belief with assigning high probability to those narratives. Sure, there are also simpler predictions we can estimate as well, but I think the thing that makes "belief in god" interesting from the perspective of judging intelligence is the complexity of the narrative. >Outside of…

To kill two birds with one stone, as soon as animals demonstrate something more akin to tribal-level human behavior of religious worship (the grand fusion of self, environment, social, and mystery into the realm of higher powers that must be feared) and not postmodernist definitions of worship (the grand disconnection of the self from all other factors to appeal to an idealized self that can never manifest thank you CIA involvement in the postmodernist art movement and their unique ability to capture scary mommies everywhere) then I'll consider the idea that animals are demonstrating religious behavior.

And further more, if animals DO have religious behavior, then artificial intelligence research across the entire board is WOEFULLY inadequate to represent such a core part of the neurological interaction that generates religious behavior... a core behavioral capacity that was somehow missed in every single behaviorist research paper ever published since mankind mastered animal husbandry.

You either get in bed with the idea that only humans worship gods or you have to fundamentally throw all of psychology, sociology, and animal behavior studies completely out the window.

Re: Researchers reach human parity in conversational speech recognition

#158
post #109
post #91

I'll admit I'm not very interested in speech recognition of this nature when it can't disambiguate the speaker. ie. the way Amazon Echo and other voice recognition systems can't tell the difference between a human in the room and the TV. Even when one might be clearly a female voice vs. a male. None of the voice recognition systems on the market learn my voice distinctly from my wife's or sons, and I don't want their…

for what it's worth, the "OK Google" functionality in my Android phone is trained against my voice and does a pretty good job at rejecting my wife's and coworkers commands.

For what it's worth, even after some training none of the phones in my office seem to pick up their owner's voice rather than someone else telling their own phone to run a search (we're all mid 30s male).

My phone rarely listens to me until I hold down the home button but one guy I know, who has a slow, deep voice, triggers Google to start listening in normal conversation all the time.

Re: Researchers reach human parity in conversational speech recognition

#159

Earlier quoted context omitted.

I've always evaluated audio codecs on differential signals - subtract A from A'. For telephony codecs, there are formal tests for MOS and/or PESQ .

Which, isn't substantially different than evaluating a codec on introduced SNR. These classic signal processing analyses tell you nothing about the correct operation of a psychoacoustic model codec design (i.e. MP3, AAC, Vorbis, etc.) An analysis of A (original signal) vs. A' (signal passed through a compression-decompression cycle) is just extracting the quantization noise introduced by the codec. That provides no i…

I am not familiar with the term "introduced SNR"... Googling.

No joy.

I can't really follow your thinking above; sorry.

In audio, there is distortion, which is correlated with the signal. Noise is uncorrelated. Codec error would seem very much to me to be at least much more correlated than uncorrelated. MP3 artifacts sound at least more "phase-ey" than they sound anything like quantitization noise ( at least to me ). This may be because I have heard badly aligned tape machines make that sort of error happen. "Phase-ey" also triggers my (feeble) mind into thinking about allpass filters as the model for the error.

In terms of intelligibility, it's possible to improve intelligibility by adding noise, and by adding clipping - aviation comms does this at times. What destroys intelligibility is phenomes being destroyed by phase changes and bad amplitide errors ( where there are actually good amplitude errors ).

I've run an ABACUS voice quality analyzer several times, and I don't think it's purely a distortion analyzer - adding clipping at least can improve MOS/PESQ score surprisingly. Even more surprisingly, there's no mechanism available for calibrating gain staging on one.

Re: Researchers reach human parity in conversational speech recognition

#160

Earlier quoted context omitted.

Which, isn't substantially different than evaluating a codec on introduced SNR. These classic signal processing analyses tell you nothing about the correct operation of a psychoacoustic model codec design (i.e. MP3, AAC, Vorbis, etc.) An analysis of A (original signal) vs. A' (signal passed through a compression-decompression cycle) is just extracting the quantization noise introduced by the codec. That provides no i…

I am not familiar with the term "introduced SNR"... Googling. No joy. I can't really follow your thinking above; sorry. In audio, there is distortion, which is correlated with the signal. Noise is uncorrelated. Codec error would seem very much to me to be at least much more correlated than uncorrelated. MP3 artifacts sound at least more "phase-ey" than they sound anything like quantitization noise ( at least to me ).…

[deleted]
Post reply on HN