Live data from Hacker News

DeepSpeech 0.6

hacks.mozilla.org

31–40 of 68 posts

Re: DeepSpeech 0.6

#31

Earlier quoted context omitted.

Hi! I'm from the RWTH team ( https://paperswithcode.com/sota/speech-recognition-on-libris... ). Our best system (2.3% WER) (trained only on the Librispeech data, i.e. much less data than your model) currently is a hybrid HMM/NN model, and you are right, the acoustic model uses a BLSTM. However, we have shown in other work ( https://www-i6.informatik.rwth-aachen.de/publications/downlo... ) that you can expect similar…

Hi! > I guess you prefer an "end-to-end" model over a hybrid HMM/NN model, for simplicity, right? Simplicity and ease of targeting other languages, yes. We're a small team. > As far as I remember, you use CTC, right? I always wondered why you have chosen CTC, and not some better model, like RNN-T, RNA, or some of the streaming attention variants. We started DeepSpeech in 2016, before these recent developments for end…

> Nice! Can you share a link to the streaming attention code?

It uses our TensorFlow framework Returnn (https://github.com/rwth-i6/returnn). We currently only have some configs/code online for hard attention variants, or segmental models. They can be found here: https://github.com/rwth-i6/returnn-experiments/tree/master/2...

Note that the configs are maybe not so cleaned up, as this is very much research. This is for a paper we submitted to ICASSP. We did not publish the paper yet elsewhere, but I can send you a copy by mail (just contact me: albzey@gmail.com).

Also, as this is research focused, the encoder here is also a BLSTM, because we wanted to compare this work to other global soft attention models, and have the comparison mostly focused on the streaming attention modeling aspect. And also there is some lookahead, which is currently unlimited. So it would need a few modifications to really be used for online streaming. I'm also not sure whether this is the best model, or whether some of the many other variants (MoChA, RNN-T, etc) are maybe better.

Edit: I forgot, we also have some local windowed attention variants, which can also be applied for streaming: https://github.com/rwth-i6/returnn-experiments/tree/master/2...

Re: DeepSpeech 0.6

#32
post #19

I do not understand how to use Deepspeech even in the most simple use case. 1. I want to teach it ten words. How do I do this? 2. I want to speak into my microphone (available as a Pulseaudio device) and recognise the words and output the words as a text stream on stdout. How do I do this? This is the documentation: https://deepspeech.readthedocs.io/en/v0.6.0/Python-Examples.... https://deepspeech.readthedocs.io/en/v…

I think you're being a little harsh. 1. Train it using the training dataset but change all word labels that aren't in your 10 word list to "OTHER_WORD" or whatever. There's probably not much point doing this. Alternatively I guess you can restrict the beam search to only look for your 10 words. Yes this is really complicated - it's a speech recognition engine - they are really complicated! 2. The API docs look pretty…

> Huh well that's a flaw the neglected to mention! I guess you can't really do what you want yet.

The documentation is outdated. That is no longer true, intermediateDecode is cheap. Thanks for noticing, I'll fix it.

Re: DeepSpeech 0.6

#33
post #19

I do not understand how to use Deepspeech even in the most simple use case. 1. I want to teach it ten words. How do I do this? 2. I want to speak into my microphone (available as a Pulseaudio device) and recognise the words and output the words as a text stream on stdout. How do I do this? This is the documentation: https://deepspeech.readthedocs.io/en/v0.6.0/Python-Examples.... https://deepspeech.readthedocs.io/en/v…

So, funny story: I wanted #2 and did it myself. I was similarly frustrated with the lack of documentation.

It's been a little while since I got it running, but I basically got a siri clone working. If you want to test it out, I can try to answer questions / whatever problems pop up.

The code is here: https://github.com/shawwn/DeepSpeech/commit/01f5cf8d39c356ae...

As far as I know, you can simply run speech_to_text.sh. It will connect to your microphone and start dumping out transcribed audio to stdout.

It wasn't super easy, but once you spend a little time with the code you can sort of figure out ways to get it to do what you want.

EDIT: By the way, people will try to convince you that you're nuts and that the documentation is crystal clear and so on. Know this: It's not just you. I had the exact same experience. It's a recurring theme in AI programming.

The only thing to do is to either find someone else who feels similarly, or roll up your sleeves and dive into the code.

Re: DeepSpeech 0.6

#34
post #5

> It achieves a 7.5% word error rate on the LibriSpeech test clean benchmark Anyone have a comparison for how good/bad that is compared to other solutions, and what it means for practical usage, if that can be guessed at from a single number?

Author here. I would add to what the sibling comments have mentioned by saying that SotA results should be taken with a grain of salt. Our engine is capable of streaming (processing the audio as it's being recorded), which is not doable with architectures that have bidirectional decoders or attention mechanisms that require the whole encoder input ahead of time. For real world applications, this is absolutely crucial…

NVIDIA has QuartzNet which contains only 19M weights and achieves 3.9% on test clean without language model and less then 3 with LM. Code (Pytorch): https://github.com/NVIDIA/NeMo Paper: https://arxiv.org/pdf/1910.10261.pdf

Re: DeepSpeech 0.6

#35
post #19

I do not understand how to use Deepspeech even in the most simple use case. 1. I want to teach it ten words. How do I do this? 2. I want to speak into my microphone (available as a Pulseaudio device) and recognise the words and output the words as a text stream on stdout. How do I do this? This is the documentation: https://deepspeech.readthedocs.io/en/v0.6.0/Python-Examples.... https://deepspeech.readthedocs.io/en/v…

So, funny story: I wanted #2 and did it myself. I was similarly frustrated with the lack of documentation. It's been a little while since I got it running, but I basically got a siri clone working. If you want to test it out, I can try to answer questions / whatever problems pop up. The code is here: https://github.com/shawwn/DeepSpeech/commit/01f5cf8d39c356ae... As far as I know, you can simply run speech_to_text.sh…

>By the way, people will try to convince you that you're nuts and that the documentation is crystal clear and so on

No it really is simple: the original implementation had a base of prefabulated amulition grammar, this was surmounted by a malleable logarithmic function in such a way that the two main spurving vocabularies were in a direct line with the panametric grammar fields. The latter simply consists of marzlevanes fitted to the ambifacient morpheme waneshaft to eliminate side utterances. This main winding is of the normal lotus-o-deltoid type, just placed in panendermic semi-boloid stators. Basically every second conductor is connected by a nonreversible tremmie pipe to the differential grammeters.

All these words are in: https://en.wikipedia.org/wiki/Turboencabulator

Re: DeepSpeech 0.6

#36
This reminds me of something I would love to see happen but I don't have the skills to put it all together. I really think there's some potential merit to a reading coach app(lication) that listens to someone read and looks for weaknesses/disorders/etc compared to a trained model. It could provide those diagnostics to an educator, guide the content to focus on those, coach the reader directly, etc.

It all seems very doable based on what I see in the technology today, I just don't have the skills to do it.

Re: DeepSpeech 0.6

#38

Earlier quoted context omitted.

Author here. I would add to what the sibling comments have mentioned by saying that SotA results should be taken with a grain of salt. Our engine is capable of streaming (processing the audio as it's being recorded), which is not doable with architectures that have bidirectional decoders or attention mechanisms that require the whole encoder input ahead of time. For real world applications, this is absolutely crucial…

Hi! I'm from the RWTH team ( https://paperswithcode.com/sota/speech-recognition-on-libris... ). Our best system (2.3% WER) (trained only on the Librispeech data, i.e. much less data than your model) currently is a hybrid HMM/NN model, and you are right, the acoustic model uses a BLSTM. However, we have shown in other work ( https://www-i6.informatik.rwth-aachen.de/publications/downlo... ) that you can expect similar…

You 100% tuned your hyper parameters and LM for librespeech, and only using librispeech data probably helped the result. When you start training systems for LVCSR it's normal for the WER to go up on some test sets. (This is why you can never expect a commercial system to get SOTA on librispeech, despite it being much better system in reality)

Re: DeepSpeech 0.6

#39
post #5

> It achieves a 7.5% word error rate on the LibriSpeech test clean benchmark Anyone have a comparison for how good/bad that is compared to other solutions, and what it means for practical usage, if that can be guessed at from a single number?

7.5% WER is much much worse than the state of the art, but it's also important to recognize that LibriSpeech is kind of a ridiculous data set. It's 1000 hours of volunteers reading public domain novels out loud. It's very unlikely to match your use case. It's widely used in academia purely because it's larger than other freely available data sets.

Re: DeepSpeech 0.6

#40

Earlier quoted context omitted.

Hi! I'm from the RWTH team ( https://paperswithcode.com/sota/speech-recognition-on-libris... ). Our best system (2.3% WER) (trained only on the Librispeech data, i.e. much less data than your model) currently is a hybrid HMM/NN model, and you are right, the acoustic model uses a BLSTM. However, we have shown in other work ( https://www-i6.informatik.rwth-aachen.de/publications/downlo... ) that you can expect similar…

You 100% tuned your hyper parameters and LM for librespeech, and only using librispeech data probably helped the result. When you start training systems for LVCSR it's normal for the WER to go up on some test sets. (This is why you can never expect a commercial system to get SOTA on librispeech, despite it being much better system in reality)

But these results are about in line with other CTC systems. The original Baidu paper [1] had 7.89% with original Deep Speech and 5.33% with "Deep Speech 2" where they threw more complicated networks at it and added another 11K hours of external speech traingin data.

[1] https://arxiv.org/pdf/1512.02595v1.pdf

Post reply on HN