I'll believe it when I actually see it. I'm a native of a reasonably small language spoken by about a million people and never have I ever seen a good automatic translation for it. The only translations that are good are the ones that have been manually entered, and those that match the structure of the manually entered ones. I think the sentiment is laudable and wish godspeed to the people working on this, but for t…
No Language Left Behind
31–40 of 166 posts
Re: No Language Left Behind
#32What is a "low resource language"?
hey there, I work on this project. We categorize a language as low-resource if there are fewer than 1M publicly available, de-duplicated bitext samples. also see section 3, table 1 in the paper: https://research.facebook.com/publications/no-language-left-...
Re: No Language Left Behind
#33Jeff Wang here with my fellow Meta AI colleague Angela Fan from No Languages left Behind, seeing the comments flowing through. If you want to ask us anything, go for it!
Re: No Language Left Behind
#34My concern with this is that in low resource languages the unavoidable biases of the ML models might overpower their own organic development. We shrug off all the little quirks of machine translated text because it usually gets the point across, and we recognize them as quirks because most of what we read was written by real people with no such quirks. But when most of what you read contain those quirks, I fear those…
Re: No Language Left Behind
#35I'm not entirely sure why low resource languages are seen as such a high priority for AI research. It seems that by definition there's little payoff to solving translation for them.
Re: No Language Left Behind
#36I see the mixture model is ~ 300 GB and was trained on 256 GPUs.
I assume distilled versions can easily be run on one GPU.
Re: No Language Left Behind
#37Earlier quoted context omitted.
hey there, I work on this project. We categorize a language as low-resource if there are fewer than 1M publicly available, de-duplicated bitext samples. also see section 3, table 1 in the paper: https://research.facebook.com/publications/no-language-left-...
hey, this sounds silly but I can't seem to find a link of all the languages covered in the 200 hundred languages. I've looked at the website and the blogpost and neither have a readily available link. Seems like a major oversight. There is of course a drop down in both but the languages there are a lot less than 200. I'm particularly interested in a list of the 55 African languages for example.
Re: No Language Left Behind
#38Jeff Wang here with my fellow Meta AI colleague Angela Fan from No Languages left Behind, seeing the comments flowing through. If you want to ask us anything, go for it!
What is the greatest insight you gained and could share with non-experts from working on this project?
Re: No Language Left Behind
#39What are hardware requirements to run this? I see the mixture model is ~ 300 GB and was trained on 256 GPUs. I assume distilled versions can easily be run on one GPU.
Re: No Language Left Behind
#40My concern with this is that in low resource languages the unavoidable biases of the ML models might overpower their own organic development. We shrug off all the little quirks of machine translated text because it usually gets the point across, and we recognize them as quirks because most of what we read was written by real people with no such quirks. But when most of what you read contain those quirks, I fear those…
In a worst case you can end up with the Scots Wikipedia situation, where some power editor created a bunch of pages using an entirely fabricated, overly stereotypical language and that influenced what people thought Scots actually was.