Live data from Hacker News

High-performance speech recognition with no supervision at all

ai.facebook.com

61–66 of 66 posts

Re: High-performance speech recognition with no supervision at all

#61
post #43

Will Facebook license this tech out? I can't be the only one who's noticed Google's speech recognition has gotten significantly less accurate the last couple of years, probably the result of some cost-cutting strategy.

There was a time a few years back where Google's speech recognition was virtually flawless (at least for my voice) but nowadays it verges on useless. I've had similar experiences with Google Maps (specifically navigation) and Google Translate. My understanding is that the implementation of these products is now mostly ML driven whereas in previous iterations ML was just one aspect of it. Curious to know how much trut…

Indeed. Growing up playing with voice recognition on Windows I knew how to talk to computers. You speak clearly, enunciate your consonants, and keep a consistent pace and even tone. When I used my "computer voice" on Android I could carry on a text or IM conversation with my phone sitting in my pocket. Nowadays it hasn't gotten any better at understanding my natural drawl but the "computer voice" fills the sentences with bizarre punctuation and randomly-capitalized words that make me look like a lunatic.

I really don't know just what the hell my phone wants from me anymore.

Re: High-performance speech recognition with no supervision at all

#62

Earlier quoted context omitted.

People keep downvoting me, because "we need you to send a proof", but sorry I don't want to share Snowden's fate, but I know for fact IOS/Android, most TVs and major TV cable boxes (USA) are indeed analyzing your speech and sending back home _keywords_ from your conversations. I can imagine that's how they avoid being sued into oblivion for clear 4A violation, as metatags have been considered not a violation of your…

Ok, for those that want proof, its pretty simple to do. 1) we know that sending voice data to "HQ" costs power 2) we know that live transcription costs a huge wedge of power 3) we know that wakeword matching is quite power efficient. (see https://rhasspy.readthedocs.io/en/latest/wake-word/ , https://github.com/MycroftAI/mycroft-precise ) So, in a quite room we know that to save power and data, devices won't be stream…

Did you try the example I described at your home?

You not going to distinguish what people want to BUY from Google search as good as from conversation. When I google "Ferrari" it may mean I am looking for Ferrari wallpaper, Ferrari stats, Ferrari parts, or want to buy new Ferrari. When I have a conversation with someone about buying Ferrari, the conclusion IS I am ready to buy a Ferrari.

Re: High-performance speech recognition with no supervision at all

#63

Question: are there any efforts to communally create / crowd source training material for neural networks? I’m thinking language like this, but also labelled imagery for, for example, face detection, which works better on white people. Has anyone attempted to create a way for people to create and donate labelled data to a dataset?

Common Voice ( https://commonvoice.mozilla.org/ ) is one such project. I've urged the people I know who have less-common accents to contribute, but I'm not sure they have.

Good, high quality, wide coverage, labeled datasets are expensive to assemble. Most companies don't want to give them away. You can find a number from academia, though.

Re: High-performance speech recognition with no supervision at all

#65

Earlier quoted context omitted.

Babelfish like technology would be a dream come true for me. At this point in time, google translate speech recognition works well when the speaker speaks a bit slow and uses simple sentences. The translation is the same, works well for simple sentences even though the translations in my language sound weird because it uses the equivalent of old English. There is still a way to go before speech recognition can recogn…

> There is still a way to go before speech recognition can recognize fast speech with colloquialisms, and probably even a longer way to go before it can translate longer sentences with abstractions into something that is clear and concise in the target language. Did you watch the video tweeted by Facebook's CTO? https://twitter.com/schrep/status/1395766932104572928 The speed, accent, and word choices of the speaker w…

I wonder if this or similar tech can be used for animals.

Re: High-performance speech recognition with no supervision at all

#66
post #54
post #52

Earlier quoted context omitted.

Are you sure? A translation error played part in the dropping of the atomic bomb on Hiroshima. https://pangeanic.com/knowledge/the-worst-translation-mistak... Language has great ambiguity, depends heavily on time period, context and tone. https://www.reddit.com/r/funny/comments/2chfge/you_need_some... That could lead to problems even without translation.

So maybe don't use the tech yet in such high-stakes situations. In my personal life, an automated translation error is very unlikely to cause a serious problem, while an automated driving error could kill me. (Fwiw, if you read the comments at your second link, you'll find the image is a fake.)

Same for autonmous driving. Don't use it on the highway but in the city. Lower speed, lower risk of dying. Fake or not, doesn't make my point invalid.
Post reply on HN