A few weeks ago, a deep learning researcher at one of the world's leading speech groups told me off-the-record that offline, human-parity speech recognition would be "coming soon" to mobile devices. Not sure s/he realized just how soon that would be. Even though state-of-the-art ASR is really expensive to train, recognition is extremely cheap to run, even on lower-power devices. [1][2] With specialized silicon, you c…
Speech-to-Text-WaveNet: End-to-end sentence level English speech recognition
11–20 of 43 posts
Re: Speech-to-Text-WaveNet: End-to-end sentence level English speech recognition
#12Does anyone on HN do active research in this field? Could I pick your brain for a survey of the best papers (especially review papers) on the subject?
Re: Speech-to-Text-WaveNet: End-to-end sentence level English speech recognition
#13This seems super useful for most speech recognition - understanding context. It doesn't seem like the mainstream engines (Alexa, Google Voice, Siri) are context aware. Why not?
This is what I'm solving at Optik. Helping you manage the things that you care about in the place that you are, and NOT exposing your personal details to cloud computation.
Re: Speech-to-Text-WaveNet: End-to-end sentence level English speech recognition
#14A few weeks ago, a deep learning researcher at one of the world's leading speech groups told me off-the-record that offline, human-parity speech recognition would be "coming soon" to mobile devices. Not sure s/he realized just how soon that would be. Even though state-of-the-art ASR is really expensive to train, recognition is extremely cheap to run, even on lower-power devices. [1][2] With specialized silicon, you c…
A consumer focused human parity ASR service will disrupt so many industries, including mine. I run a human powered transcription service where we transcribe files with high accuracy. I am just waiting for the day when our transcribers can work off a auto-generated transcript instead of typing it all up manually. I'll pay good money for a service where I can just send a file and get a 80-90% accurate transcript with s…
Re: Speech-to-Text-WaveNet: End-to-end sentence level English speech recognition
#15Re: Speech-to-Text-WaveNet: End-to-end sentence level English speech recognition
#16So, I'm not the only one seeing this issue. It seems like many recent AI papers want to look as impressive as possible, wile giving you as little implementation info as possible. This bothers me, because it opposes the very purpose of research publication.
Re: Speech-to-Text-WaveNet: End-to-end sentence level English speech recognition
#17Earlier quoted context omitted.
A consumer focused human parity ASR service will disrupt so many industries, including mine. I run a human powered transcription service where we transcribe files with high accuracy. I am just waiting for the day when our transcribers can work off a auto-generated transcript instead of typing it all up manually. I'll pay good money for a service where I can just send a file and get a 80-90% accurate transcript with s…
I hope you realize your business is about to go out of business. The only reason you can charge people now is because the automatic recognition sucks compared to humans.
Re: Speech-to-Text-WaveNet: End-to-end sentence level English speech recognition
#18Earlier quoted context omitted.
I hope you realize your business is about to go out of business. The only reason you can charge people now is because the automatic recognition sucks compared to humans.
We do super-human-parity transcripts. Our transcripts are insanely accurate, even for challenging files. I'm sure computers will be able to do that one day, but Singularity would have already happened by then, wiping out many businesses. I for one look forward to Singularity and hope that we will contribute to it in some way.
Re: Speech-to-Text-WaveNet: End-to-end sentence level English speech recognition
#19Earlier quoted context omitted.
I hope you realize your business is about to go out of business. The only reason you can charge people now is because the automatic recognition sucks compared to humans.
We do super-human-parity transcripts. Our transcripts are insanely accurate, even for challenging files. I'm sure computers will be able to do that one day, but Singularity would have already happened by then, wiping out many businesses. I for one look forward to Singularity and hope that we will contribute to it in some way.
Re: Speech-to-Text-WaveNet: End-to-end sentence level English speech recognition
#20Perhaps future communication applications can have a WaveNet on either end, which learns the voice of the person you're communicating with and then only sends text after a certain point in the conversation?
I'm coming at this from a point of ignorance though, so correct me if I've made erroneous assumptions.