Live data from Hacker News

No Language Left Behind

ai.facebook.com

91–100 of 166 posts

Re: No Language Left Behind

#91

Earlier quoted context omitted.

Looking at the list, I see a lack of Native American languages. Did anyone try to contact the tribes during this?

We interviewed speakers of low-resource languages from all over the world to understand the human need for this kind of technology --- what do people actually want, how would they use it, and what's the quality they would find useful? Many low-resource languages lack data online, but are spoken by millions. However, many indigenous languages are spoken by smaller numbers of people, and we are definitely interested in…

I'll take that in good faith, but I will say Facebook has been a particular pain for many tribal folks given its true name policy and banning people who it thinks are using a fake name. Yellow Horse was one that was wildly reported, but their are others. Mostly anything that takes the form Adjective Noun. Had a rather painful thread with someone claiming to be a Facebook employee that defended this practice. I haven't heard of anyone reaching out, and Lord knows we could of used the help because COVID has been a particular disaster for language preservation even with an extremely high vaccination rate.

I do admit I'm a bit bitter given another of the big silicon valley companies (Apple) claiming they specifically help the TCUs (Tribal Colleges and Universities) when I can find no one that knows about this help other than taking our money for product at the same price as other accredited educational institutions.

Re: No Language Left Behind

#92
I wonder if spy agencies have already developed, but not published, high-quality SMT methods for lots of minority and little-known languages. :-(

(Edit: and speech-to-text models.)

Re: No Language Left Behind

#93
post #60

Earlier quoted context omitted.

I think it will interesting when it runs into a language (e.g. Dakota) where the women and men speak differently. Should be an interesting test.

Doesn't seem to be a big issue for Arabic, where verbs are gendered (so in the sentence "I am going to the store", the verb "to go" will be either masculine or feminine, reflecting the speaker's gender).

> so in the sentence "I am going to the store", the verb "to go" will be either masculine or feminine, reflecting the speaker's gender

But there the rules are the same for everyone. This is not true in general; there are languages where men and women speak according to different rules.

Here's a selection from Empires of the Word:

> These works [written by women] are usually written in Emesal, 'the fine tongue', a separate dialect of Sumerian, well documented in scribal dictionaries. In dialogue works this dialect is used for the speech of goddesses. It differs from standard Sumerian, Emegir, 'the princely tongue', both in vocabulary (including the names of many gods) and also in pronunciation (consonants by and large being articulated farther forward in the mouth); it differs not at all in its grammar. For example, when the goddess Inanna is affecting to repel the advances of an importunate suitor, she cries:

> kuli Mulila šu bamu emeše daŋen amaŋu lulaše ta munaben amaŋu Gašangale lulaše ta munaben

> Friend of Enlil, let me free! Let me go to my house! What lie shall I tell my mother? What lie shall I tell my mother Ningal?

> Both Enlil and Ningal are, of course, gods. In Emegir this would have been (with the differences highlighted):

> kuli Enlila šu bamu eŋuše gaŋen amaŋu lulaše ana munaben amaŋu Ningale lulaše ana munaben

Re: No Language Left Behind

#94

My concern with this is that in low resource languages the unavoidable biases of the ML models might overpower their own organic development. We shrug off all the little quirks of machine translated text because it usually gets the point across, and we recognize them as quirks because most of what we read was written by real people with no such quirks. But when most of what you read contain those quirks, I fear those…

Won't people trying to learn a low resource language as as a second language also bring their influence?

Re: No Language Left Behind

#95

I'll believe it when I actually see it. I'm a native of a reasonably small language spoken by about a million people and never have I ever seen a good automatic translation for it. The only translations that are good are the ones that have been manually entered, and those that match the structure of the manually entered ones. I think the sentiment is laudable and wish godspeed to the people working on this, but for t…

> I ever seen a good automatic translation for it. > > When Google Translate regularly struggles even with big pairs such as German-English-German, I have reservations about someone making it work for languages where datasets are orders of magnitude smaller.

I speak a language where I've never seen any translation for it... and when translated manually, my mum totally butchers the meaning lol.

Either way, any work in this area is more than welcome, but damn it's a hard problem.

Re: No Language Left Behind

#97
post #2

I'm not entirely sure why low resource languages are seen as such a high priority for AI research. It seems that by definition there's little payoff to solving translation for them.

I don't really remember the exact numbers anymore, but covering only the top 5 languages will cover maybe 40% of the world population, while covering the top 200 languages (many of them low resource) will cover maybe 90% of the world population. Some numbers (but you can not exactly infer from them such accumulated numbers): https://en.wikipedia.org/wiki/List_of_languages_by_total_num... Some more numbers from here:…

It doesn't sound like you're considering that people are very often fluent in a major language in addition to their regional one?

Re: No Language Left Behind

#98
post #19

Jeff Wang here with my fellow Meta AI colleague Angela Fan from No Languages left Behind, seeing the comments flowing through. If you want to ask us anything, go for it!

Hi, I'm looking but can't seem to find instructions on how to do tokenization. Where is spm model, is it "flores200_sacrebleu_tokenizer_spm.model" or something else? And is it direct or spm -> dict? Or how to prime model for a specific language pair?

Re: No Language Left Behind

#99
post #98
post #19

Jeff Wang here with my fellow Meta AI colleague Angela Fan from No Languages left Behind, seeing the comments flowing through. If you want to ask us anything, go for it!

Hi, I'm looking but can't seem to find instructions on how to do tokenization. Where is spm model, is it "flores200_sacrebleu_tokenizer_spm.model" or something else? And is it direct or spm -> dict? Or how to prime model for a specific language pair?

We tokenize with the flores-200 spm model, correct. To generate from the model, check out the instructions here: https://github.com/facebookresearch/fairseq/tree/nllb/exampl...

Re: No Language Left Behind

#100

Earlier quoted context omitted.

We have a full list here (copy pastable): https://github.com/facebookresearch/flores/tree/main/flores2... and Table 1 of our paper ( https://research.facebook.com/publications/no-language-left-... ) has a complete list as well.

Nice to see Esperanto made the cut — the only artificial language to do so, AFAICT.

I was happy to see that as well!
Post reply on HN