Live data from Hacker News

Live-caption glasses let deaf people read conversations [video]

youtube.com

171–180 of 181 posts

Re: Live-caption glasses let deaf people read conversations [video]

#171
post #170

Earlier quoted context omitted.

Single-mic source separation is possible in an unsupervised manner today that could probably work better than beamforming both compute-wise and with regard to implementation difficulty (you'd just need a lot of recordings to represent the space of sounds you want to separate).

Do you have any references for this or a link to a commercial service? I'm currently in the process or trying to extract some background voices in a video (an interview where the faint background conversation is in English and the loud overdub is in Bulgarian). I tried Melodyne but it seems to only separate on pitch, not volume, and the pitchs are too similar (mono, three voices, all female) and words are made of lot…

https://arxiv.org/abs/2110.10739

I haven't seen it provided as a commercial service or free model yet, but there is open source code for Mixit that lets you train using the open source / canned FUSS dataset.

Re: Live-caption glasses let deaf people read conversations [video]

#172

Earlier quoted context omitted.

Single-mic source separation is possible in an unsupervised manner today that could probably work better than beamforming both compute-wise and with regard to implementation difficulty (you'd just need a lot of recordings to represent the space of sounds you want to separate).

Maybe a combination? Even simple beamforming/stereo would be helpful to help display the speaker's location. For example, the "speaker 1" tag could appear on the left, center, or right of the display to give a spatial clue where they are located.

You could only display the speaker's location if you had some way to associate streams with the individuals. So you'd have to train an audio-visual association model.

The thing you could do is train a localizer on the separated audio. Phase is estimated by the source separation process, so you can actually train an ML model provided you have some ground truth (e.g. estimated human locations from camera detections)

Re: Live-caption glasses let deaf people read conversations [video]

#173
post #170

Earlier quoted context omitted.

Do you have any references for this or a link to a commercial service? I'm currently in the process or trying to extract some background voices in a video (an interview where the faint background conversation is in English and the loud overdub is in Bulgarian). I tried Melodyne but it seems to only separate on pitch, not volume, and the pitchs are too similar (mono, three voices, all female) and words are made of lot…

https://arxiv.org/abs/2110.10739 I haven't seen it provided as a commercial service or free model yet, but there is open source code for Mixit that lets you train using the open source / canned FUSS dataset.

Thanks! That led me to this which looks like a good place to start:

https://github.com/google-research/sound-separation/blob/mas...

Re: Live-caption glasses let deaf people read conversations [video]

#174

Earlier quoted context omitted.

Thank you for looking at XRAI Glass! 1. For multiple simultaneous speakers of comparable volume, it’s only as good as the underlying speech-to-text engines we’ve implemented/integrated, which is currently not very good. It’s active area of research and engineering for us and we believe we’ll make strides to improve things; but, as you rightly point out, solving the crosstalk problem is very difficult. For the more ge…

Seems like the multiple microphone beamforming source separation algorithms are getting pretty good these days, maybe just adding a lot more mics would help? Could you have an AI model that extracts some characteristics of the speaker's voice for each individual word, then translates that to color and font? If the model was not confident about a word it could show slightly blurred, if it was loud it could be bold, pe…

Regarding additions mics, yes that enabled more advanced spatial processing, especially when arranged in a determined and calibrated geometry. One could imagine glasses with several mics placed at optimal locations on the frames. Such multichannel audio could then be processed into multichannel spatial audio streams.

As for the rest, thank you for wonderful ideas! Everything you propose is technically possible. The difficulties arise first in assessing the increased benefit to users versus the increased complexity of user experience, and second in prioritizing the work versus other features. Over time we do hope to add additional selectable “skins”, which is to say different UI designs, that allow users to choose UI’s from a wide range. Everything from the simple to complex layouts, from accessible to exotic color palettes, from professional to playful themes, etc. I could definitely see more advanced visual representations of transcription uncertainty showing up in such optional skins.

Re: Live-caption glasses let deaf people read conversations [video]

#175

Earlier quoted context omitted.

Thank you for looking at XRAI Glass! 1. For multiple simultaneous speakers of comparable volume, it’s only as good as the underlying speech-to-text engines we’ve implemented/integrated, which is currently not very good. It’s active area of research and engineering for us and we believe we’ll make strides to improve things; but, as you rightly point out, solving the crosstalk problem is very difficult. For the more ge…

Wow! Thanks for the response. This is exciting work, and I’m pleased to see it being iterated on. Re: #2. I’m assuming the varying depth is manually-controlled? Or is it automated by some method? If it’s manual, can the adjustment be made while transcription is active? In other words, can I change the focal distance to match the speaker without interrupting the speaker? All in all, cool stuff! Best of luck with the w…

Thank you for the praise and encouragement!

Indeed, as Dan said, one can change the depth on the fly. In fact, one goal of the development team is to make as much functionality as possible changeable on the fly. For example, you can currently change subtitle depth, pinning, and size on the fly, spoken language, subtitle language, microphones, and audio settings on the fly, etc.

I’d love for every setting and feature to support on the fly changes. That said, some things are currently fixed for a session, such as recording audio, and some third party software we utilize is less dynamic and forgiving of changes on the fly. For better or worse, in our software world of today, the sage advice of The IT Crowd “Have you tried turning on and off again?” still seems to hold with pragmatic force. And it still holds with XRAI ... sometimes ;-)

Re: Live-caption glasses let deaf people read conversations [video]

#176

The number one thing these glasses/software need to solve is that the words match the speech for a one to one conversation in a quiet environment. Eg doctors visit. I think they are very close. We just got the Nreal/Xrai setup a few days ago for deaf from birth wife (hearing husband) She grew up lipreading but integrated more with signing and deaf community as an adult. She has a cochlear implant but can not understa…

Thank you for this feedback! Btw, we'll support better adjustment of the subtitle position very soon. We did just add many additional font size options as well. If you haven't already, please consider joining our Discord server to provide feedback at any time: https://discord.gg/7HjyDJ3JAz

Re: Live-caption glasses let deaf people read conversations [video]

#177

Earlier quoted context omitted.

ASL is not based on french either. It's related to LSF, the sign language used in france, but that also isn't based on french. The modern sign languages emerged among deaf populations and have completely different grammar and morphology from the spoken languages of the cultures surrounding their origins.

Thank you for correcting me. That's a deep rabbit hole

Spoken languages are linear. Sign languages are not. The grammar is very different.

Sign languages have multiple articulators: two hands, face, eye gaze direction, shoulders, trunk. These can all work together to show multiple things at the same time. Spoken languages can really do only one thing at a time (with a few minor suprasegmentals such as tone).

You can construct a signed version of a spoken language, which may be useful for things like quoting book titles and other cases where you need to represent the exact words of a spoken language in signed form, but it's not common to use that for everyday communication, because the hands move a lot slower than the small muscles of the mouth and throat.

(Linguistics is a fascination of mine. Sign language linguistics are especially interesting.)

Re: Live-caption glasses let deaf people read conversations [video]

#178
post #3

Earlier quoted context omitted.

Folks that use ASL use fingerspelling which is of course just written english, no? https://www.lifeprint.com/asl101/fingerspelling/fingerspelli...

Nope. https://www.signingsavvy.com/article/45/The+difference+betwe... Here's a fun example. ASL allows, maybe even requires, negation after the statement. An interpreter friend of mine was interpreting Wayne's World in a mixed crowd. The whole " ... NOT!" joke gets laughs from the hearing audience and the Deaf audience doesn't understand why.

I think that could be interpreted. Statement + NO is the standard word order in ASL, but there would usually be a suprasegmental element. That is, the negation is also (or, sometimes, only) shown with a headshake which spreads over the entire length of the statement. Leaving out the suprasegmental, making it a flat statement, and then pausing before the NO might perhaps work. Maybe.

However, ASL also makes much, much heavier use of rhetorical questions than English does. You might even introduce yourself with "MY NAME WHAT? [NAME]" (i.e., "What is my name? [Name]"). So perhaps it would just look like you're doing that.

(Disclaimer: I don't know ASL. I know some Irish Sign Language, which is related, but dropped out before completing my interpreter training. I have a bit of a fascination with sign language linguistics, but I'm no expert.)

Re: Live-caption glasses let deaf people read conversations [video]

#179
post #28

I'm a hearing person and I've spent a summer interning in a 50/50 mixed Deaf and hearing research group. My take is that this is a huge UI improvement for AI speech to text, which a lot of Deaf people are already using to listen to conversations. It seems particularity great because it allows this technology to provide situational awareness while, for example, walking. It's important to remember, though that, for con…

Why can’t they also make glasses that translate sign language into English audio or text then?

Translating sign languages to spoken languages is really, really hard. Sign languages have a lot of features that don't really show up in spoken languages.

Let's take an example: the sign for "send email" in ASL (http://www.lifeprint.com/asl101/pages-signs/e/email.htm). If I point at you at the end of the sign it could mean "I will send you an email." If I point at myself it could mean "You should send me an email" or "Did you send me an email" depending on my facial expression. If I point off into space it could mean "I'm sending an email." If I start by pointing at you and then end the sign by pointing off into space it could mean "You should send an email." So your translator AI needs not only to understand the facial expressions and movements of the signer, but also the spacial relationships of everyone in the conversation. And that is just one aspect of the difficulty - there are many other features of sign languages that are just as hard to translate.

Perhaps this is the sort of thing that future AI systems could do. But it is quite complex.

Re: Live-caption glasses let deaf people read conversations [video]

#180
post #154

Earlier quoted context omitted.

Which group is that? Disabled people, or their allies, or someone else?

Many countries have aging populations, so the proportion of people who have a personal reason to care about issues like sight degeneration is going to increase significantly over time.

bingo
Post reply on HN