Or, if it had OCR capabilities, you could just hold a sheet of paper in front of yourself and say "what's this?" and it would explain the text to you.
Live-caption glasses let deaf people read conversations [video]
151–160 of 181 posts
Re: Live-caption glasses let deaf people read conversations [video]
#152Re: Live-caption glasses let deaf people read conversations [video]
#153Background: researched this space for a graduate degree. There are a few issues that are unanswered by this video (which isn't intended to be a technical deep dive, but I don't see any related links in the video description): 1. How do these glasses handle multiple simultaneous speakers? Based on the display I saw, it shows the speakers' words sequentially, which starts to fall apart in real-world environments, espec…
The last time I worked on anything here, there were a number of problems with transmitting highly compressed, low resolution video data. The consumer devices could not handle just sending the packets. In my project we were annotating real time video and only sending back the annotations, but even that would cause devices to overheat and the applications to fail in really interesting ways.
Re: Live-caption glasses let deaf people read conversations [video]
#154Earlier quoted context omitted.
There is a huge population group that I am hoping will demand and make accessibility much more refined. Everything from the size of text to lighting levels to subtitles
Which group is that? Disabled people, or their allies, or someone else?
Re: Live-caption glasses let deaf people read conversations [video]
#155Background: researched this space for a graduate degree. There are a few issues that are unanswered by this video (which isn't intended to be a technical deep dive, but I don't see any related links in the video description): 1. How do these glasses handle multiple simultaneous speakers? Based on the display I saw, it shows the speakers' words sequentially, which starts to fall apart in real-world environments, espec…
Thank you for looking at XRAI Glass! 1. For multiple simultaneous speakers of comparable volume, it’s only as good as the underlying speech-to-text engines we’ve implemented/integrated, which is currently not very good. It’s active area of research and engineering for us and we believe we’ll make strides to improve things; but, as you rightly point out, solving the crosstalk problem is very difficult. For the more ge…
If you had more microphones placed in multiple spots on the glasses, such as up to 5 microphones - 1 in the center, 2 on the frames/end pieces, and perhaps a final 2 on the arms/temples. Then that would be able to catch conversation coming at a person from behind them, the sides, or directly in front, etc.
Re: Live-caption glasses let deaf people read conversations [video]
#156Re: Live-caption glasses let deaf people read conversations [video]
#157Background: researched this space for a graduate degree. There are a few issues that are unanswered by this video (which isn't intended to be a technical deep dive, but I don't see any related links in the video description): 1. How do these glasses handle multiple simultaneous speakers? Based on the display I saw, it shows the speakers' words sequentially, which starts to fall apart in real-world environments, espec…
Thank you for looking at XRAI Glass! 1. For multiple simultaneous speakers of comparable volume, it’s only as good as the underlying speech-to-text engines we’ve implemented/integrated, which is currently not very good. It’s active area of research and engineering for us and we believe we’ll make strides to improve things; but, as you rightly point out, solving the crosstalk problem is very difficult. For the more ge…
This would be such a cool use case for the latest ChatGPT tech when it gets faster in the near future.
Re: Live-caption glasses let deaf people read conversations [video]
#158Background: researched this space for a graduate degree. There are a few issues that are unanswered by this video (which isn't intended to be a technical deep dive, but I don't see any related links in the video description): 1. How do these glasses handle multiple simultaneous speakers? Based on the display I saw, it shows the speakers' words sequentially, which starts to fall apart in real-world environments, espec…
Is multiple simultaneous speakers at the same volume/distance actually an important problem to solve? I already can't have a conversation if that's happening and my hearing is fine.
Re: Live-caption glasses let deaf people read conversations [video]
#159Earlier quoted context omitted.
>3. As mentioned by another commenter, this is a useful idea for people who lose their hearing later in life. That said, this is less (although certainly still) useful for people who have congenital hearing loss and primarily communicate via ASL.> Someone primarily communicates with ASL and then there's me that doesn't know ASL. I can speak to them, and they can read what I've spoken. That works pretty well. They com…
ASL is heavily inspired by English, so it's usually very easy for people to become basically conversant in it rather quickly. For signs you don't know, there's literal "finger spelling" that's part of the language, so conversational learning is greatly aided by this. Which.. aside from that, you could do this anyways. When I first started living with a deaf person, I just wrote things down on paper, and they wrote ba…
Re: Live-caption glasses let deaf people read conversations [video]
#160Background: researched this space for a graduate degree. There are a few issues that are unanswered by this video (which isn't intended to be a technical deep dive, but I don't see any related links in the video description): 1. How do these glasses handle multiple simultaneous speakers? Based on the display I saw, it shows the speakers' words sequentially, which starts to fall apart in real-world environments, espec…
Thank you for looking at XRAI Glass! 1. For multiple simultaneous speakers of comparable volume, it’s only as good as the underlying speech-to-text engines we’ve implemented/integrated, which is currently not very good. It’s active area of research and engineering for us and we believe we’ll make strides to improve things; but, as you rightly point out, solving the crosstalk problem is very difficult. For the more ge…
Could you have an AI model that extracts some characteristics of the speaker's voice for each individual word, then translates that to color and font?
If the model was not confident about a word it could show slightly blurred, if it was loud it could be bold, perhaps(Although there's some stereotype issues) you could use different fonts for different pitches, whispers could be grey, quiet could be transparent.
Maybe there's a language model that can pick up overlapping words if you don't have the constraint of needing to sort them out into who said it, just show all the possibile words that could have been said by anyone stacked together, in a "not sure" color, and maybe the wearer would eventually learn to figure it out without much effort?
You could also try to stay consistent so the same speaker gets the same colors I'd possible, and also not reuse colors for new speakers that have been recently used, to best make use of the limited bits of data in font and color.
Maybe just by showing all the words from every speaker all together like that, the wearer would be able to figure it out even if it made mistakes in the speaker identification?