Live data from Hacker News

Researchers reach human parity in conversational speech recognition

blogs.microsoft.com

161–164 of 164 posts

Re: Researchers reach human parity in conversational speech recognition

#161

Earlier quoted context omitted.

Which, isn't substantially different than evaluating a codec on introduced SNR. These classic signal processing analyses tell you nothing about the correct operation of a psychoacoustic model codec design (i.e. MP3, AAC, Vorbis, etc.) An analysis of A (original signal) vs. A' (signal passed through a compression-decompression cycle) is just extracting the quantization noise introduced by the codec. That provides no i…

I am not familiar with the term "introduced SNR"... Googling. No joy. I can't really follow your thinking above; sorry. In audio, there is distortion, which is correlated with the signal. Noise is uncorrelated. Codec error would seem very much to me to be at least much more correlated than uncorrelated. MP3 artifacts sound at least more "phase-ey" than they sound anything like quantitization noise ( at least to me ).…

Ah, that's a mistake, that should have read "introduced [quantization] noise", which reduces the SNR in A' vs A.

If I understand you correctly, your discussion around intelligibility primarily applies to signal processing of voice vs general audio. Codecs such as MP3, AAC, etc. don't have the luxury of making assumptions about the signal content, and so aren't designed along those principles. E.g. speech codecs can generally run well at much lower bitrates than general audio codecs because they operate on a constrained domain of audio (i.e. speech).

Regarding distortion management, see Rate-distortion optimization[1] for lossy audio codecs: where the purpose is to manage distortion within the limits of the bit rate supported by a communication channel or storage medium.

[1] https://en.wikipedia.org/wiki/Quantization_(signal_processin...

Re: Researchers reach human parity in conversational speech recognition

#162

Earlier quoted context omitted.

I am not familiar with the term "introduced SNR"... Googling. No joy. I can't really follow your thinking above; sorry. In audio, there is distortion, which is correlated with the signal. Noise is uncorrelated. Codec error would seem very much to me to be at least much more correlated than uncorrelated. MP3 artifacts sound at least more "phase-ey" than they sound anything like quantitization noise ( at least to me ).…

Ah, that's a mistake, that should have read "introduced [quantization] noise", which reduces the SNR in A' vs A. If I understand you correctly, your discussion around intelligibility primarily applies to signal processing of voice vs general audio . Codecs such as MP3, AAC, etc. don't have the luxury of making assumptions about the signal content, and so aren't designed along those principles. E.g. speech codecs can…

Thanks for clarifying.

Re: Researchers reach human parity in conversational speech recognition

#163

Earlier quoted context omitted.

Do you alone in the world know what animals are thinking?

I provided the very best and most recent research for the absurd position that animals supposedly have their own animal gods, attend regular communal worshipping ceremonies, engage in animal religious ritual, and demonstrate religious behavior outside of a priori instinct and social fitnesses. You provided... passive aggressiveness. And if animals now engage in religious ritual, doesn't that make them stupid for not…

> I provided the very best and most recent research for the absurd position that animals supposedly have their own animal gods

You did nothing of the sort. You provided a link to a video of an anthropologist describing chimpanzee behaviour and indulging in a little light speculation about why they're doing what they're doing.

Nobody is claiming that animals believe in god(s), rather I'm disputing your bizarre assertion that you know that they don't. There is no way to tell.

Re: Researchers reach human parity in conversational speech recognition

#164

How long do you think it is until captioning companies / TRSs such as Captel downsize significantly because of tech like this?

Supply generates its own demand; by making captioning even cheaper, it can increase the demand for transcription services and people to check it over. There are a lot of podcasts and YT videos that could benefit from transcriptions but it's too expensive now.
Post reply on HN