Live data from Hacker News

Deep Learning enables hearing aid wearers to pick out a voice in a crowded room

spectrum.ieee.org

41–50 of 60 posts

Re: Deep Learning enables hearing aid wearers to pick out a voice in a crowded room

#41
post #7

This approach surprised me. Why are they doing feature extraction and then feeding that into a DNN? It seems much more straightforward to have the input of the network be noisy samples and the output be clean samples a la super resolution[0] in images. They probably wouldn't want to use fully-connected layers in that instance, but I don't see any fundamental barriers if they have enough computational power to run a n…

That might work, although I think there are two limitations: 1) Hearing aids have a 10ms latency budget. So no matter how much processing they can do, they're limited by how many samples they can look ahead and that limits the design of the filters. The brain can presumably look ahead further to separate sound streams so I think it's pretty impressive that ideal binary masking works. 2) Hearing aids have a power budg…

The latency and power issues can probably be fixed, assuming a good end-to-end model, by using model distillation into a wide shallow net using low-precision or even binary operations. I don't know if that would be enough - we've seen multiple order of magnitude decreases in compute requirements (think about style transfer going from hours on top-end Titan GPUs to realtime on mobile phones) but the usual target is mobile smartphones which at least have a GPU, while it seems unlikely any hearing aids will have GPUs anytime soon... I suppose a good enough squashed low-precision model could be turned into an ASIC.

Re: Deep Learning enables hearing aid wearers to pick out a voice in a crowded room

#43
I can't pick out a voice in a crowded room, or indeed separate speech from any sort of background noise. In ideal listening conditions I miss words making sentences not make sense, and often I don't realize someone has started talking until I've already missed the first sentence. However I don't have any physical hearing problem. Each time I've gotten my hearing tested I've been told my tonal hearing is perfectly normal. Yet the problems I have with picking out and understanding speech are absolutely debilitating, and I can't get anyone to understand that it is a disability and that it is real.

At my insistence my audiologist administered a speech processing test, but I was nonplussed to discover this test is completely unrealistic and did not at all match the situations I have trouble with. The way it worked was that it would mix a perfectly clear speech track with white noise, or a repetitive loop of background speech or cafeteria noise. But since the sound streams were mixed together so artificially, my brain could separate the audio streams based on source track, words or no. And since the "interruption" loops were repetitive, my brain could learn the pattern and discount it. So of course I passed that test, too. The speech processing problems I have occur in real environments when executing functions of daily life.

In the end the audiologist told me that maybe my problem is that I have ADHD and that my attention isn't able to lock-on or stay with a conversation well. He didn't know of anyone in my area who treated adults with ADHD, but promised to send me a referral. I'm guessing he never found anyone, because that referral never came. However it eventually led to me getting diagnosed and treated for ADHD on my own. (Although it took almost 2 years to even get an appointment.) I've found that getting a diagnosis and medication for ADHD has improved my life immensely. However it has not helped with the original problem; I still can't separate speech from other noises.

I resonate with the commenter who says he thinks his undiagnosed (physical) hearing loss once contributed to him losing a job. At work, I find excuses to hide/disconnect my phone because I have so much trouble making out what people are saying over a phone. I use chat and IM, and write everything down or ask for things written down. Still, sometimes I'll miss or not understand some verbal instruction and get in trouble. It also causes relationship problems - so many misunderstandings, misheard words, doing the opposite of what my spouse asked or not realizing she said something to me. I avoid some social activities because I know that background noise there will prevent me from participating, or because mishearing people might lead to a social gaffe or a dangerous misinterpretation of safety instructions.

I can't read lips to get by, either - whatever it is in my brain that affects speech processing affects lip reading equally, if not worse, and sometimes when I'm receiving the "all circuits are down" message from my speech centers, I can't even understand someone's sentence no matter how many times they repeat it. But if they write it down on a note I can understand it. In a way, it's like the inverse of dyslexia.

Anyway, not that I have much hope of an answer, but anyone know where I can go to talk about it or what kind of doctor would be actually interested and not just brush this off?

Re: Deep Learning enables hearing aid wearers to pick out a voice in a crowded room

#44
post #16
post #6

> The greatest frustration among potential users is that a hearing aid cannot distinguish between, for example, a voice and the sound of a passing car if those sounds occur at the same time. The device cranks up the volume on both, creating an incoherent din. It may be a simplification of the article that I'm misinterpreting, but as someone who got a hearing aid in early 2016, that's not how (modern) hearing aids wor…

>Hearing aids have changed my quality of life (at age 40). I've had mine for 6 months. I'm 63 now. They have changed my life as well. I still have some tinniness but the tech has been adjusting the curve and other factors each time I visit her and this last time a few weeks ago the sound quality got much better. There are still places where they don't work or get overwhelmed by background noise, like at a live baskeb…

The deeply robust elder Deaf community itself is proof that the brain tissue loss described in the article is probably not due to the hearing loss itself. I have never heard of a culturally Deaf individual experiencing this kind of degeneration of the brain. If I had to hazard a hypothesis it would be on the social end of things, such as social isolation causing decreased function, similar to to the social contexts that encourage addiction.

In other words, when one loses hearing, the culture around them fails to accommodate that leading to increased anxiety, stress, and other kinds of undesirable outcomes that are shown to impact health. It is immeasurably tragic in many ways that this phenomenon is being used to sell hearing aids and more anxiety.

Re: Deep Learning enables hearing aid wearers to pick out a voice in a crowded room

#45
post #12

Earlier quoted context omitted.

As someone with a cochlear implant who lives with the consequences of overly clever programmers who thought they'd "help" by filtering out noise and volume and whatever else... I really wish they wouldn't. This is a technology that makes me so angry some days that I sometimes wonder if it was worth getting implanted, even though I know it was.

This is something I do wonder about, in this context. I don't have a CI myself, but my 6 year old son does, and I am somewhat concerned that he is-or-might be experiencing partial sound "blindness" (meaning: sure speech processing is adequate but there are surely some things that are processed away). I have a fair amount of experience in music/sound-recording environments and it makes me somewhat sad for him that he'…

Interesting idea on using ML to help non-signers understand sign language. In this thread's context, the ML is designed to help people hear better. In visual contexts (which sign language lives in) would this hypothetical ML help low vision or blind people see better?

Because people with 20/20 vision just need a sign language dictionary handy and some patience.

Re: Deep Learning enables hearing aid wearers to pick out a voice in a crowded room

#46

Earlier quoted context omitted.

> That blows my mind ... a bit like how do you know the color green is green. Maybe it's purple, but you have been told by someone else that it's green. You may be interested in the "the map is not the territory" idea. "Green" is not a property of an object but of an observer, though "emits light at wavelength N" is a property of an object.

"Green" is not a property of an object but of an observer, But "Green" happens to be a property of other large families of objects -- especially animate objects (foliage, certain insects, birds, and fish) and, more rarely, certain inanimate but nonetheless "special" objects in nature (features in geology; the sky at certain times; and of course, rainbows). So in that sense -- while "Green" by itself doesn't seem to h…

I don't really understand your point. "Green" is a label we apply to things which fall into a certain category: namely those which under ordinary circumstances emit or reflect light of a certain wavelength. The observer-dependent things here are:

- the definition of the category (which varies depending on what "ordinary circumstances" are for the observer: people draw the boundaries of light-wavelengths differently), and

- making the judgement "this object is in/is absent from the Green category" for any given object (since our information is imperfect, and [for instance] we may only ever see an object while it is bathed in blue light).

My post was mainly intended to say "there's no paradox involved if you experience green objects differently to me: the word 'green' corresponds to a quale which isn't an inherent property of things in the universe, but an artefact of our experience". Additionally, in this comment, I point out that 'green' can indicate different qualia to different people anyway.

Re: Deep Learning enables hearing aid wearers to pick out a voice in a crowded room

#47
post #35

Earlier quoted context omitted.

I have the Advanced Bionics Harmony BTE. Since my implant is AB, I wouldn't be able to get the Nucleus Freedom 6. I have an in-ear mic, which does wonders for reducing surrounding noises and also for letting me use a phone normally, but my main issue is with the software itself; I've had issues with it since implantation and they've always been pooh-poohed by audiologists at Hopkins, Tokyo University, and Toranomon.…

Yep, that's exactly the thing I'm talking about. You'd think they could hire one deaf person at their labs to road test the things, but... On my hearing aid it's a "feature" that can be turned off. Too bad you're stuck with it.

I have had so many problems with the implant in general that are just brushed off as "well, you're unusual." It won't even stay on my head without me putting a few extra magnets on the headpiece.

Re: Deep Learning enables hearing aid wearers to pick out a voice in a crowded room

#48
post #12

Earlier quoted context omitted.

This is something I do wonder about, in this context. I don't have a CI myself, but my 6 year old son does, and I am somewhat concerned that he is-or-might be experiencing partial sound "blindness" (meaning: sure speech processing is adequate but there are surely some things that are processed away). I have a fair amount of experience in music/sound-recording environments and it makes me somewhat sad for him that he'…

Interesting idea on using ML to help non-signers understand sign language. In this thread's context, the ML is designed to help people hear better. In visual contexts (which sign language lives in) would this hypothetical ML help low vision or blind people see better? Because people with 20/20 vision just need a sign language dictionary handy and some patience.

> Interesting idea on using ML to help non-signers understand sign language.

The dream for me is something like Google Glass with an app that can subtitle spoken, written, and signed language.

> Because people with 20/20 vision just need a sign language dictionary handy and some patience.

I would think a LOT of patience... the easiest way at that point would just to have the other person fingerspell or write what they're saying; if you're watching something where that's not possible, then the dictionary will just be an exercise in frustration.

Re: Deep Learning enables hearing aid wearers to pick out a voice in a crowded room

#49
post #41

Earlier quoted context omitted.

That might work, although I think there are two limitations: 1) Hearing aids have a 10ms latency budget. So no matter how much processing they can do, they're limited by how many samples they can look ahead and that limits the design of the filters. The brain can presumably look ahead further to separate sound streams so I think it's pretty impressive that ideal binary masking works. 2) Hearing aids have a power budg…

The latency and power issues can probably be fixed, assuming a good end-to-end model, by using model distillation into a wide shallow net using low-precision or even binary operations. I don't know if that would be enough - we've seen multiple order of magnitude decreases in compute requirements (think about style transfer going from hours on top-end Titan GPUs to realtime on mobile phones) but the usual target is mo…

Not to detract from your larger point but AFAIK the style transfer thing is different. If you're willing to hardcode the style into the net you can go realtime, but the original style transfer paper is able to do different styles without retraining. So they're different algorithms. Unless the SOTA has changed recently.

Re: Deep Learning enables hearing aid wearers to pick out a voice in a crowded room

#50

Earlier quoted context omitted.

Interesting idea on using ML to help non-signers understand sign language. In this thread's context, the ML is designed to help people hear better. In visual contexts (which sign language lives in) would this hypothetical ML help low vision or blind people see better? Because people with 20/20 vision just need a sign language dictionary handy and some patience.

> Interesting idea on using ML to help non-signers understand sign language. The dream for me is something like Google Glass with an app that can subtitle spoken, written, and signed language. > Because people with 20/20 vision just need a sign language dictionary handy and some patience. I would think a LOT of patience... the easiest way at that point would just to have the other person fingerspell or write what the…

> and some patience.

Yeah, well that would be one way of handling it, but unfortunately the real world has terrible issues with not impeding my progress on that front. Not that I'm anti-learning, at all, but - personally - I'm fighting a losing battle against learning German, Swiss-German and Swiss-German Sign-language whilst also being a walking-talking-english-lesson :D

Taking the slow way, with dictionary in hand, is as you point out, an exercise in frustration (especially if the talker/signer is 6 years old).

Yes, I share your dream of something google-glass-like that can add subtitles. There are people working on this (mostly in the UAE, if memory serves). Interesting times ahead - hopefully I won't have to wait long, otherwise I'll have to do it myself and that really would take a while ;)

Post reply on HN