Live data from Hacker News

Weak supervision to isolate sign language communicators in crowded news videos

vrroom.github.io

41–50 of 53 posts

Re: Weak supervision to isolate sign language communicators in crowded news videos

#41

Earlier quoted context omitted.

I think you think it's a magic box. There's not actually such thing as a "strong language model", not in the way you're using the concept. > We hope that what we are building might at least improve the state-of-the-art there. Do you have any theoretical arguments for how and why it would improve it? If not, my concern is that you're just sucking the air out of the room. (Research into "throw a large language model at…

“Throw an LM at it” is the only approach that has ever produced human level machine translation. For theory on how a strong target-language-side LM can improve translation, even in the extreme scenario where no parallel “texts” are available, https://proceedings.neurips.cc/paper_files/paper/2023/file/7...

You're mixing up cause and effect. The transformer architecture was invented for machine translation – and it's pretty good at it! (Very far from human-level, but still mostly comprehensible, and a significant improvement over the state-of-the-art at time of first publication.) But we shouldn't treat this as anything more than "special-purpose ML architecture achieves decent results".

The GPT architecture, using transformers to do iterated predictive text, is a modern version of the Markov bot. It's truly awful at translation, when "prompted" to do so. (Perhaps surprisingly so, until you step back, look at the training data, and look at the information flow: the conditional probability of the next token isn't mostly coming from the source text.)

I haven't read that paper yet, but it looks interesting. From the abstract, it looks like one of those perfectly-valid papers that laypeople think is making a stronger claim than it is. This paragraph supports that:

> Note that these models are not intended to accurately capture natural language. Rather, they illustrate how our theory can be used to study the effect of language similarity and complexity on data requirements for UMT.

Re: Weak supervision to isolate sign language communicators in crowded news videos

#42
Using news broadcast as a training model to populate LLM is a poor precedence.

Repetition of a sign usually indicates an additional emphasis.

The dialect needs to be all covered and multiply mapped to its word.

Furthermore, YouTube has an excellent collection of really bad or fake ASL interpreters in many news broadcasts, so bad, really really bad, worse than Al Gore Hanging Chad news broadcast or the "hard-of-hearing" inset box during Saturday Night Live News broadcast.

You still need an RID-certified or CDI-certified ASL interpreter to vet the source.

https://m.youtube.com/watch?v=GwSh0dAaqIA

https://rid.org/certification/available-certifications/

Re: Weak supervision to isolate sign language communicators in crowded news videos

#43
post #13

Earlier quoted context omitted.

> It's more of a bad and broken transliteration that if you struggle to think about you can parse out and understand. it seems to be more common to see sign language interpreters now. is it just virtue signaling to have that instead of just closed captions?

Also, in India, many hearing-impaired people know only ISL.

Just so you know, "hearing-impaired" implies that a person has a flaw whether a person is born with it (it is natural to them) or impacted in later life (hearing-challenged).

Most non-offensive way to refer to a group of people without perfect hearing is "hard-of-hearing or deaf".

Re: Weak supervision to isolate sign language communicators in crowded news videos

#44
post #23

Earlier quoted context omitted.

I know an interpreter who is a CODA. Her first language was sign language, which I think helps a lot. I once asked her if she thought in English or ASL and she said ASL. During the pandemic she’d get very frustrated by the ASL she saw on the news. Her mom and deaf friends couldn’t understand them. It wasn’t long before she was on the news regularly to make sure better information was going out. She kept getting COVID…

> I once asked her if she thought in English or ASL and she said ASL It's common to think that we think in languages, but at a fundamental level, we simply don't. Ever have the experience of not being able to remember the word for something? you know exactly what you are thinking, but can't come up with the word. If you thought in your language, this wouldn't happen.

The other day I forgot the name of a street in my town, right as I was trying to tell someone to turn onto it.

I said, "I don't know the name of the street - It's the son of (CHARACTER) from (VIDEO GAME)" and then I remembered the name. I was completely right about the game character, apparently I'd remembered that mnemonic but temporarily forgotten the character's actual name.

Re: Weak supervision to isolate sign language communicators in crowded news videos

#45

Earlier quoted context omitted.

I know an interpreter who is a CODA. Her first language was sign language, which I think helps a lot. I once asked her if she thought in English or ASL and she said ASL. During the pandemic she’d get very frustrated by the ASL she saw on the news. Her mom and deaf friends couldn’t understand them. It wasn’t long before she was on the news regularly to make sure better information was going out. She kept getting COVID…

I got a insanely dumb question, that probably has an obvious answer: why is it so critically important to have ASL? It seems to me you could skip having someone to the side adding the layers of emoting imperfectly and just watch the anchor and captions? Note, I'm definitely the one missing something. When first met with a contradiction, check your premises, is my motto, not lecture on why they're wrong.

Most sign language users are functionally illiterate.

https://academic.oup.com/jdsde/article/17/1/1/359085

Re: Weak supervision to isolate sign language communicators in crowded news videos

#46
post #39

Earlier quoted context omitted.

I know an interpreter who is a CODA. Her first language was sign language, which I think helps a lot. I once asked her if she thought in English or ASL and she said ASL. During the pandemic she’d get very frustrated by the ASL she saw on the news. Her mom and deaf friends couldn’t understand them. It wasn’t long before she was on the news regularly to make sure better information was going out. She kept getting COVID…

During Covid, my son (who is deaf and attends a deaf school) has issued masks with transparent windows at the front, especially for assisting with lip-reading for deaf users. This is in Switzerland though - I don't know if this innovation reached across the Atlantic ;)

They did, but it wasn't like they were available in March of 2020 when the world was halting. When getting news and communication was important.

Re: Weak supervision to isolate sign language communicators in crowded news videos

#47
1st: I sign ASL not ISL like the OP is talking about.

In the ASL world, most news translations into ASL are delayed or sped up from the person talking and/or the captions if they happen to also be available.

You are going to have sync problems.

Secondly, it's not just moving the hands, body movements, facial expressions, etc all count in ASL , I'm betting they count in ISL as well.

Thirdly the quality of interpretation can be really bad. Horrendous. it's not so common these days, but it was fairly common that speakers would hire an interpreter and mistakenly hire someone willing to just move their arms randomly. I had it happen once at a doctors office. The "interpreter" was just lost in space. The doctor and I started writing things down and the interpreter seemed a little embarrassed at least.

Sometimes they hire sign language students, you can imagine hiring a first year french student to interpret for you, it's no different really. Sometimes they mean well, sometimes they are just there for the paycheck.

I bet it's a lot worse with ISL, because it's still very new, most students are not taught in ISL, there are only about 300 registered interpreters for millions of deaf people in India. https://islrtc.nic.in/history-0

We are still very much struggling with vocal to English transcriptions using AI. Despite loads of work from lots of companies and researchers. They are getting better, and in ideal scenarios are actually quite useful. Unfortunately the world is far from ideal.

The other day on a meeting with 2 people using the same phone. The AI transcription was highly confused and it went very, very wrong.

I'm not trying to discourage you, and it's great to see people trying. I wish you lots of success, just know it's not an easy thing and I imagine lots of lifetimes of work will be needed to generate useful signed language to written language services that are on-par with the best of the voice to text systems we have today.

Re: Weak supervision to isolate sign language communicators in crowded news videos

#48

Earlier quoted context omitted.

What captions? There’s no widely-used written form of ASL. If you meant the English-language captions, well, it should be apparent why some people prefer content to be dubbed in their native language rather than reading subtitles in a different language that they understand less well.

Yes I meant $NATIVE_LANGUAGE_OF_VIEWER, and that example didn't help me I'm afraid, I'm quite a dense one! I appreciated you checking if I meant written ASL captions. :) The example of dubs left me confused because dubs brings in an extra sense that isn't applicable in the ASL case. There, in either case, we're watching someone who isn't the character. I can't square that with the existential fervor the woman in OP f…

> Yes I meant $NATIVE_LANGUAGE_OF_VIEWER

Right, the native language of most ASL users isn’t English. It’s ASL. And there is no such thing as captions in ASL (because it has no widely-used written form), which is why you have interpretation.

Re: Weak supervision to isolate sign language communicators in crowded news videos

#49

Earlier quoted context omitted.

Yes I meant $NATIVE_LANGUAGE_OF_VIEWER, and that example didn't help me I'm afraid, I'm quite a dense one! I appreciated you checking if I meant written ASL captions. :) The example of dubs left me confused because dubs brings in an extra sense that isn't applicable in the ASL case. There, in either case, we're watching someone who isn't the character. I can't square that with the existential fervor the woman in OP f…

> Yes I meant $NATIVE_LANGUAGE_OF_VIEWER Right, the native language of most ASL users isn’t English. It’s ASL. And there is no such thing as captions in ASL (because it has no widely-used written form), which is why you have interpretation.

[deleted]

Re: Weak supervision to isolate sign language communicators in crowded news videos

#50
post #29

Earlier quoted context omitted.

There are layers to thought, and the layer that is most conscious is, at least for me (and from what I hear a lot of other people), in my native language. Further, achieving fluency in a second language is often associated with skipping the intermediate step of translating from English, with thoughts materializing first in the second language. You're correct that there does seem to be a layer lower than that—one that…

you can talk to yourself in your head using your native language. But that's not evidence that you are thinking in your language, it's thinking of your language. when you do math, are you thinking in your language? when you drive a car or play a video game or a sport, or make love, are you thinking in your language? I'll answer for you, no, you aren't. Do people who grow up without language (there have been plenty ex…

Thinking is probably not a single unified process, but several simultaneous activities at different levels. Including the unconscious ones. So you might very well both be right, as we may be thinking sometimes simultaneously in language and in other ways (like when we talk and cook, when we write, etc.)
Post reply on HN