Live data from Hacker News

Weak supervision to isolate sign language communicators in crowded news videos

vrroom.github.io

21–30 of 53 posts

Re: Weak supervision to isolate sign language communicators in crowded news videos

#21
post #9

Earlier quoted context omitted.

Thanks for the feedback. You raise great points and this was the reason why we wrote this post, so that we can hear from people where the actual problem lies. On a related note, this sort of explains why our model is struggling to fit on 500 hours of our current dataset (even on the training set). Even so, the current state of automatic translation for Indian Sign Language is that, in-the-wild, even individual words…

I think you think it's a magic box. There's not actually such thing as a "strong language model", not in the way you're using the concept. > We hope that what we are building might at least improve the state-of-the-art there. Do you have any theoretical arguments for how and why it would improve it? If not, my concern is that you're just sucking the air out of the room. (Research into "throw a large language model at…

“Throw an LM at it” is the only approach that has ever produced human level machine translation.

For theory on how a strong target-language-side LM can improve translation, even in the extreme scenario where no parallel “texts” are available, https://proceedings.neurips.cc/paper_files/paper/2023/file/7...

Re: Weak supervision to isolate sign language communicators in crowded news videos

#22

> I believe that we can solve continuous sign language translation convincingly American Sign Language is not English, in fact, it's not even particularly close to English. Much of the language is conveyed with body movements outside of the hands and fingers, particularly with facial expressions and "named placeholders." > All this is to say, that we need to build a 5000 hour scale dataset for Sign Language Translati…

I know an interpreter who is a CODA. Her first language was sign language, which I think helps a lot. I once asked her if she thought in English or ASL and she said ASL. During the pandemic she’d get very frustrated by the ASL she saw on the news. Her mom and deaf friends couldn’t understand them. It wasn’t long before she was on the news regularly to make sure better information was going out. She kept getting COVID…

I got a insanely dumb question, that probably has an obvious answer: why is it so critically important to have ASL? It seems to me you could skip having someone to the side adding the layers of emoting imperfectly and just watch the anchor and captions? Note, I'm definitely the one missing something. When first met with a contradiction, check your premises, is my motto, not lecture on why they're wrong.

Re: Weak supervision to isolate sign language communicators in crowded news videos

#23

> I believe that we can solve continuous sign language translation convincingly American Sign Language is not English, in fact, it's not even particularly close to English. Much of the language is conveyed with body movements outside of the hands and fingers, particularly with facial expressions and "named placeholders." > All this is to say, that we need to build a 5000 hour scale dataset for Sign Language Translati…

I know an interpreter who is a CODA. Her first language was sign language, which I think helps a lot. I once asked her if she thought in English or ASL and she said ASL. During the pandemic she’d get very frustrated by the ASL she saw on the news. Her mom and deaf friends couldn’t understand them. It wasn’t long before she was on the news regularly to make sure better information was going out. She kept getting COVID…

>I once asked her if she thought in English or ASL and she said ASL

It's common to think that we think in languages, but at a fundamental level, we simply don't. Ever have the experience of not being able to remember the word for something? you know exactly what you are thinking, but can't come up with the word. If you thought in your language, this wouldn't happen.

Re: Weak supervision to isolate sign language communicators in crowded news videos

#24

Earlier quoted context omitted.

I know an interpreter who is a CODA. Her first language was sign language, which I think helps a lot. I once asked her if she thought in English or ASL and she said ASL. During the pandemic she’d get very frustrated by the ASL she saw on the news. Her mom and deaf friends couldn’t understand them. It wasn’t long before she was on the news regularly to make sure better information was going out. She kept getting COVID…

I got a insanely dumb question, that probably has an obvious answer: why is it so critically important to have ASL? It seems to me you could skip having someone to the side adding the layers of emoting imperfectly and just watch the anchor and captions? Note, I'm definitely the one missing something. When first met with a contradiction, check your premises, is my motto, not lecture on why they're wrong.

Think of it more like "why is it important to have Spanish as an option for captions in an area that knows it has a large audience of L1 Spanish speakers?". English/ is pretty much a foreign language to many signers. Their native language is ASL or whatever other sign language they know, and these languages aren't just 1:1 mappings of words in the local dominant language to hand signals. They have many dimensions of expression for encoding meaning like facial expression/body motions, speed of the sign, amount of times they repeat the sign that have grammatical meaning that spoken languages express with inflection (noun cases, different verb forms) or additional words. For example, where an English speaker would use words like "very" or "extremely" or choose adjectives with more intense connotations an ASL signer would repeat or exaggerate the sign they want to emphasize (often by signing it more quickly, but it frequently involves intensifying multiple parts of the sign like the entire motion or the facial expression as well).

Re: Weak supervision to isolate sign language communicators in crowded news videos

#25

Earlier quoted context omitted.

I know an interpreter who is a CODA. Her first language was sign language, which I think helps a lot. I once asked her if she thought in English or ASL and she said ASL. During the pandemic she’d get very frustrated by the ASL she saw on the news. Her mom and deaf friends couldn’t understand them. It wasn’t long before she was on the news regularly to make sure better information was going out. She kept getting COVID…

I got a insanely dumb question, that probably has an obvious answer: why is it so critically important to have ASL? It seems to me you could skip having someone to the side adding the layers of emoting imperfectly and just watch the anchor and captions? Note, I'm definitely the one missing something. When first met with a contradiction, check your premises, is my motto, not lecture on why they're wrong.

What captions? There’s no widely-used written form of ASL.

If you meant the English-language captions, well, it should be apparent why some people prefer content to be dubbed in their native language rather than reading subtitles in a different language that they understand less well.

Re: Weak supervision to isolate sign language communicators in crowded news videos

#26
post #23

Earlier quoted context omitted.

I know an interpreter who is a CODA. Her first language was sign language, which I think helps a lot. I once asked her if she thought in English or ASL and she said ASL. During the pandemic she’d get very frustrated by the ASL she saw on the news. Her mom and deaf friends couldn’t understand them. It wasn’t long before she was on the news regularly to make sure better information was going out. She kept getting COVID…

> I once asked her if she thought in English or ASL and she said ASL It's common to think that we think in languages, but at a fundamental level, we simply don't. Ever have the experience of not being able to remember the word for something? you know exactly what you are thinking, but can't come up with the word. If you thought in your language, this wouldn't happen.

It's often said that a picture is worth a thousand words. And there's no succinct picture that reliably convey that concept, at the same time. Human cognition must be simply cross-modal.

Re: Weak supervision to isolate sign language communicators in crowded news videos

#27
post #23

Earlier quoted context omitted.

I know an interpreter who is a CODA. Her first language was sign language, which I think helps a lot. I once asked her if she thought in English or ASL and she said ASL. During the pandemic she’d get very frustrated by the ASL she saw on the news. Her mom and deaf friends couldn’t understand them. It wasn’t long before she was on the news regularly to make sure better information was going out. She kept getting COVID…

> I once asked her if she thought in English or ASL and she said ASL It's common to think that we think in languages, but at a fundamental level, we simply don't. Ever have the experience of not being able to remember the word for something? you know exactly what you are thinking, but can't come up with the word. If you thought in your language, this wouldn't happen.

There are layers to thought, and the layer that is most conscious is, at least for me (and from what I hear a lot of other people), in my native language. Further, achieving fluency in a second language is often associated with skipping the intermediate step of translating from English, with thoughts materializing first in the second language.

You're correct that there does seem to be a layer lower than that—one that can materialize as either a native or a second language—but it's not inaccurate to talk about which language we "think" in because many of us actually do constantly materialize thoughts as language without any intention of speaking them.

Re: Weak supervision to isolate sign language communicators in crowded news videos

#28

> I believe that we can solve continuous sign language translation convincingly American Sign Language is not English, in fact, it's not even particularly close to English. Much of the language is conveyed with body movements outside of the hands and fingers, particularly with facial expressions and "named placeholders." > All this is to say, that we need to build a 5000 hour scale dataset for Sign Language Translati…

> American Sign Language is not English I'm not sure I understand your point. Chinese is also not English but machine translation of Chinese to English can be done. You're right that laypeople often assume, wrongly, that a given country's sign language is an encoding of the local spoken language. In reality it's usually a totally different language in its own right. But that shouldn't mean that translation is fundame…

They didn't say it's fundamentally impossible, they said you need way more than 5000 hours to do it and that you'd need better quality sources than news interpreters.

Re: Weak supervision to isolate sign language communicators in crowded news videos

#29
post #23

Earlier quoted context omitted.

> I once asked her if she thought in English or ASL and she said ASL It's common to think that we think in languages, but at a fundamental level, we simply don't. Ever have the experience of not being able to remember the word for something? you know exactly what you are thinking, but can't come up with the word. If you thought in your language, this wouldn't happen.

There are layers to thought, and the layer that is most conscious is, at least for me (and from what I hear a lot of other people), in my native language. Further, achieving fluency in a second language is often associated with skipping the intermediate step of translating from English, with thoughts materializing first in the second language. You're correct that there does seem to be a layer lower than that—one that…

you can talk to yourself in your head using your native language. But that's not evidence that you are thinking in your language, it's thinking of your language. when you do math, are you thinking in your language? when you drive a car or play a video game or a sport, or make love, are you thinking in your language? I'll answer for you, no, you aren't.

Do people who grow up without language (there have been plenty examples, deaf people for example) simply not think? do cats and dogs and chimpanzees not think?

Re: Weak supervision to isolate sign language communicators in crowded news videos

#30
post #9

> I believe that we can solve continuous sign language translation convincingly American Sign Language is not English, in fact, it's not even particularly close to English. Much of the language is conveyed with body movements outside of the hands and fingers, particularly with facial expressions and "named placeholders." > All this is to say, that we need to build a 5000 hour scale dataset for Sign Language Translati…

Thanks for the feedback. You raise great points and this was the reason why we wrote this post, so that we can hear from people where the actual problem lies. On a related note, this sort of explains why our model is struggling to fit on 500 hours of our current dataset (even on the training set). Even so, the current state of automatic translation for Indian Sign Language is that, in-the-wild, even individual words…

> Do you think if we make a system for bad/broken transliteration and funnel it through ChatGPT, it might give meaningful results?

No, because ChatGPT's training data has practically no way of knowing what a real sign language looks like, since there's no real written form of any sign language and ChatGPT learned its languages from writing.

Sincerely: I think it's awesome that you're taking something like this on, and even better that you're open to learning about it and correcting flawed assumptions. Others have already noted some holes in your understanding of sign, so I'll also just note that I think a solid brush up on the fundamentals of what language models are and aren't is called for—they're not linguistic fairy dust you can sprinkle on a language problem to make it fly. They're statistical machines that can predict likely results based on their training corpus, which corpus is more or less all the text on the internet.

I'm afraid I'm not in a good position to recommend beginner resources (I learned this stuff in university back before it really took off), but I've heard good things about Andrej Karpathy's YouTube channel.

Post reply on HN