Live data from Hacker News

Meta's Omnilingual MT for 1,600 Languages

ai.meta.com

41–50 of 56 posts

Re: Meta's Omnilingual MT for 1,600 Languages

#41

I’m very wary of celebrating Meta’s language work when the company was credibly found to have contributed to the genocide against the Rohingya in Myanmar, and separately, to human rights abuses against Tigrayans during the conflict in northern Ethiopia. Be careful whose sins you’re laundering. https://www.amnesty.org/en/latest/news/2025/02/meta-new-poli... https://www.amnesty.org/en/latest/news/2023/10/meta-failure-.…

I had the same reaction to this post. Mainly because one of Meta's explanations for the lapse was that they didn't have moderators who understood the local language

You hear that folks? Lack of localization kills.

Re: Meta's Omnilingual MT for 1,600 Languages

#42
post #23

I'll be looking at this in detail. I've started a company to do similar things, https://6k.ai I'm currently concentrating on better data gathering for low-resource languages. When you look in detail at data like Common Crawl, finepdfs, and fineweb, (1) they are really lacking quality data sources if you know where to look, and (2) the sources they have are not processed "finely" enough (e.g. finepdfs classify each pa…

It's sad that I didn't see any languages on your website from Australia, where there are hundreds of languages that need translating.

Re: Meta's Omnilingual MT for 1,600 Languages

#44

That's a high count, but still a bit away from "Omni". Usual count is between 4k and 8k depending the source. But the first 1k might be the hardest, certainly.

1.6k languages is for how many we were able to find more or less reliable evaluation data (mostly thanks to Bible translators and all those who contributed to BOUQuET). Out of the remaining several thousand languages, we expect the OMT models to support understanding (but not generation) for a significant proportion, due to cross-lingual generalisation between similar languages. So it’s not truly “omni” in the sense of supporting every single language on Earth, but it’s our best effort to do so, and probably the most “omni” models existing today.

Re: Meta's Omnilingual MT for 1,600 Languages

#45

Just spent a long time trying to find where you can download any of these weights. Is it open weight? If so, why isn't there just a straight link to the models?

It is not open weight as of today (unfortunately, for the reasons out of control of us the authors, we weren’t able to release the weights). All we could release is part of the evaluation data. I hope this will change in a while.

Re: Meta's Omnilingual MT for 1,600 Languages

#46
post #24

Meta released No Language Left Behind (NLLB) [1], I think in 2022. I wonder why this in not "NLLB 2.0"? These companies love introducing new names to confuse things [1] https://ai.meta.com/research/no-language-left-behind/

This project is absolutely NLLB 2.0 in spirit. However, we decided to reserve the name “OMT-NLLB” only to the subset of the new models that have encoder-decoder architecture similar to the original NLLB-200. The other models are called “OMT-LLaMA” and have classical LLM architecture. The idea here (and we had to emphasize it to justify the project internally) is that we are developing not just new models but a recipe for massive multilinguality that can be integrated into general-purpose LLMs.

Re: Meta's Omnilingual MT for 1,600 Languages

#47

That's a high count, but still a bit away from "Omni". Usual count is between 4k and 8k depending the source. But the first 1k might be the hardest, certainly.

1.6k languages is for how many we were able to find more or less reliable evaluation data (mostly thanks to Bible translators and all those who contributed to BOUQuET). Out of the remaining several thousand languages, we expect the OMT models to support understanding (but not generation) for a significant proportion, due to cross-lingual generalisation between similar languages. So it’s not truly “omni” in the sense…

So, hyperchilio-lingual would be more accurate, and myriad-lingual would be even behind all documented existing human language. But I guess marketing team is not that found of precision in philological considerations.

Re: Meta's Omnilingual MT for 1,600 Languages

#48

That's a high count, but still a bit away from "Omni". Usual count is between 4k and 8k depending the source. But the first 1k might be the hardest, certainly.

1.6k languages is for how many we were able to find more or less reliable evaluation data (mostly thanks to Bible translators and all those who contributed to BOUQuET). Out of the remaining several thousand languages, we expect the OMT models to support understanding (but not generation) for a significant proportion, due to cross-lingual generalisation between similar languages. So it’s not truly “omni” in the sense…

Is there interest in benchmarking the proprietary LLMs for translation? Curious as I often use Gemini 3 Flash, but I have no idea how good it is for my language family. I prefer open models (in fact the smaller the better for offline), but it'd be useful to know how well the Big Three do.

Re: Meta's Omnilingual MT for 1,600 Languages

#50
post #23

I'll be looking at this in detail. I've started a company to do similar things, https://6k.ai I'm currently concentrating on better data gathering for low-resource languages. When you look in detail at data like Common Crawl, finepdfs, and fineweb, (1) they are really lacking quality data sources if you know where to look, and (2) the sources they have are not processed "finely" enough (e.g. finepdfs classify each pa…

It's sad that I didn't see any languages on your website from Australia, where there are hundreds of languages that need translating.

It’s a small sample and not specifically ones we’re working on. It’s biased towards alternative scripts for visual interest.

Australian languages are definitely interesting! and I will say, from what I’ve seen, Australian government (and other orgs) have done better than most to help document them (in recent years, at least)

Post reply on HN