Live data from Hacker News

No Language Left Behind

ai.facebook.com

11–20 of 166 posts

Re: No Language Left Behind

#11
post #2

I'm not entirely sure why low resource languages are seen as such a high priority for AI research. It seems that by definition there's little payoff to solving translation for them.

The point is that there are lots of humans who speak these languages and use tech. They just don’t use Wikipedia so getting a good translation corpus going was harder.

Re: No Language Left Behind

#12
post #2

I'm not entirely sure why low resource languages are seen as such a high priority for AI research. It seems that by definition there's little payoff to solving translation for them.

I don't really remember the exact numbers anymore, but covering only the top 5 languages will cover maybe 40% of the world population, while covering the top 200 languages (many of them low resource) will cover maybe 90% of the world population.

Some numbers (but you can not exactly infer from them such accumulated numbers): https://en.wikipedia.org/wiki/List_of_languages_by_total_num...

Some more numbers from here: https://www.sciencedirect.com/science/article/pii/S016763931...

"96% of the world’s languages are spoken by only 4% of its people."

Although this statement is more about the tail from the approx 7000 languages.

Re: No Language Left Behind

#13
So they have a system that can translate to languages for which there isn't as much data as English, Spanish, etc. Waiting for a Twitter thread from a native speaker of one of these "low resource languages" to let us know how good the actual translations are. Cynically, I'd venture that they hired some native speakers to cherry pick their best translations for the story books. But mostly this just seems like a nice bit of PR (calling it a "breakthrough", etc.). I can't imagine this is going to help anyone who actually speaks a random, e.g., Nilo-Saharan language.

Re: No Language Left Behind

#14
post #2

I'm not entirely sure why low resource languages are seen as such a high priority for AI research. It seems that by definition there's little payoff to solving translation for them.

The examples given are, with native speaker numbers, Assamese (15 million), Catalan (4 million) and Kinyarwanda (10 million). These alone are more than an Australia.

Furthermore, Facebook considers the internet to consist of Facebook and Wikipedia (Zero).

I view this as just another extension of their Next Billion initiative, an effort to ensure that another billion people are monopolised by Facebook.

That's the payoff.

Re: No Language Left Behind

#15
post #11
post #2

I'm not entirely sure why low resource languages are seen as such a high priority for AI research. It seems that by definition there's little payoff to solving translation for them.

The point is that there are lots of humans who speak these languages and use tech. They just don’t use Wikipedia so getting a good translation corpus going was harder.

And it's both cumulative across all those languages (see above), cheap/amortized (if you can do a good multilingual NMT for 50 languages, how hard can 50+1 languages be?), and many of those languages are likely to grow both in terms of sheer population and in GDP. (Think about South Asian or African countries like Indonesia or Nigeria.) The question isn't why are FB & Google investing so much in powerful multilingual models which handle hundreds of languages, but why aren't other entities as well?

Re: No Language Left Behind

#16
The analogy I like the most is that they've found the "shape" of languages in high dimensions, and if you rotate the shape for English the right way, you get an unreasonably good fit for the shape of Spanish, again for all the other languages.

We're at a point where it's now possible to determine the shape of every language, provided there are enough speakers of the language left who are both able and willing to help.

Once done, Facebook can then commodify their dissent, and sell it back to them in their native language.

Re: No Language Left Behind

#17
post #6
post #4

Earlier quoted context omitted.

I think the reason low resource languages are prioritized is to compensate for the fact that AI research normally has a tendency to marginalize these languages.

yes, but what principles justify the importance placed on low resource languages?

Low resource in this context means that there are few resources available to train a neural network with, not that there are few speakers. Although many low resource languages have relatively few speakers, there are also ones with tens of millions of speakers.

The reason for emphasis is in my opinion twofold: 1) Allowing these people to use the fancy language technology in their own language is good in and of itself. 2) Training neural networks on fewer resources is more difficult than using more resources and therefore a fun and interesting challenge.

Re: No Language Left Behind

#18

So they have a system that can translate to languages for which there isn't as much data as English, Spanish, etc. Waiting for a Twitter thread from a native speaker of one of these "low resource languages" to let us know how good the actual translations are. Cynically, I'd venture that they hired some native speakers to cherry pick their best translations for the story books. But mostly this just seems like a nice b…

Twitter may not be representative imho because of the short text. It should first come to a problem of reliable language detection, and Twitter is quite often wrong there

Re: No Language Left Behind

#19
Jeff Wang here with my fellow Meta AI colleague Angela Fan from No Languages left Behind, seeing the comments flowing through. If you want to ask us anything, go for it!

Re: No Language Left Behind

#20
post #19

Jeff Wang here with my fellow Meta AI colleague Angela Fan from No Languages left Behind, seeing the comments flowing through. If you want to ask us anything, go for it!

Are all the 200x200 translations going directly or is English (or another language) used as an intermediate for some of them?
Post reply on HN