Live data from Hacker News

Microsoft Research make breakthrough in audio speech recognition

blogs.technet.com

11–20 of 46 posts

Re: Microsoft Research make breakthrough in audio speech recognition

#11
post #8

Vlingo, Siri, and others have been doing speaker independent auto-adapting speech recognition for years and talking about systems requiring 'training' and improvements there sound like this article is 5 years old. Great to see innovation in this space but this article is very light on detail.

Vlingo has LITERALLY never gotten anything I said right, ever. Just a data point.

Re: Microsoft Research make breakthrough in audio speech recognition

#13

Imagine the power of this for students. This would have made school so much easier. Simply record every lecture and then use this to search for keywords. Awesome.

On a different note, imagine the power of this for DRM control and censorship.

But impressive and very useful.

Re: Microsoft Research make breakthrough in audio speech recognition

#14
post #8

Vlingo, Siri, and others have been doing speaker independent auto-adapting speech recognition for years and talking about systems requiring 'training' and improvements there sound like this article is 5 years old. Great to see innovation in this space but this article is very light on detail.

It is my understanding (albeit based on limited knowledge) that Siri, like other Nuance-powered systems that make a call to the server, are actually "trained" continuously by the huge amount of sample speech they receive by real users.

The true "breakthrough" here would be if Microsoft made a voice recognition system that could run entirely on a device (no internet connection needed) and accurately understand speech without terabytes of training data or a local user training session. I can't tell from the article if this is what Microsoft is claiming.

Also, it appears that "Deep Neural Network" isn't the most common term of art here. DNN appears to be a synonym for "Deep Belief Network".[1] Can anyone confirm?

[1] http://www.scholarpedia.org/article/Deep_belief_networks

Re: Microsoft Research make breakthrough in audio speech recognition

#15
This seems very related to this http://www.youtube.com/watch?v=ZmNOAtZIgIk speak by Andrew Ng. It is a 40min speak, but he explains very simply how all this works for images and some examples about the audio case. It is incredible how using this deep learning techniques we can teach this "neural networks" to recognize such complicated patterns. It is like reverse engineering the brain's algorithms.

BTW I took his Coursera's course about Machine Learning and it was great! I also recommend it A LOT to gather basic ML knowledge.

Re: Microsoft Research make breakthrough in audio speech recognition

#18
post #16

For those keeping score, google's image feature extractor shares the same core principles as microsoft's speech recognizer. EDIT: by keeping score I mean keeping track of which techniques are being used where.

Am I the only one who gets tired of people keeping score like this? Can't we just accept that many of the large companies are seriously innovative?

(Sorry, I know I'm being cranky)

Re: Microsoft Research make breakthrough in audio speech recognition

#19
post #5

The demo site ( http://www.msravs.com/audiosearch_demo/ ) blocks browsers other than IE and Firefox based on the user agent string. Use WebKit's developer tools to change your user agent and you'll be able to get in.

Why alienate such a large segment of users after pouring so much money into their technology? The web is getting weirder.

If a company invests in multiple markets, they should be prepared to do well in some markets and badly in others. Bing isn't as good as Google. Android isn't as well-designed as Metro. Yes, Android stole Apple's market, and, yes, Apple stole someone else's market. The large technology companies are deadlocked on multiple fronts. That fuels fierce competition and inspires excellence and choice. However, companies should accept they just aren't the best at everything. Let us make our own choices based on what's best for us.

Re: Microsoft Research make breakthrough in audio speech recognition

#20

Imagine the power of this for students. This would have made school so much easier. Simply record every lecture and then use this to search for keywords. Awesome.

On a different note, imagine the power of this for DRM control and censorship. But impressive and very useful.

What? What do speech recognition and fingerprinting have in common? I don't see how this research applies for DRM...

Censorship, maybe. And even then, you can't filter conversations in real-time, only maybe 'flag' people with forbidden words.

Post reply on HN