Live data from Hacker News

No Language Left Behind

ai.facebook.com

111–120 of 166 posts

Re: No Language Left Behind

#111
post #19

Jeff Wang here with my fellow Meta AI colleague Angela Fan from No Languages left Behind, seeing the comments flowing through. If you want to ask us anything, go for it!

If your goal is to make inclusive translation more widely available why license the models under a non-commercial license? This basically makes it impossible to use legally (or at least without a lot of legal risk) for essentially anyone due to the vague definition of what's commercial. Is Facebook hurting for money and looking to commercially license this model on request?

Re: No Language Left Behind

#112
>REAL-WORLD APPLICATION

>Translating Wikipedia for everyone

Hmmm.

While there is very definitely utility in doing things like this, I do kinda fear "poisoning the well"-like effects of feeding (even partially-) AI-generated-data into extremely common AI-data-sources.

There's some info on it in a blog post[1] and the MediaWiki "Content translation" page[2], but does anyone know of any studies on the quality of the translations produced? I can absolutely see it being a huge time-saver for people who are essentially fluent in both (there's a lot of semi-mechanical drudgery in translating stuff like this that could be mostly eliminated)... but people are pretty darn good at choosing the easy option of trusting whatever they're given rather than being as careful as they should be. It kinda feels like it runs the risk of passively encouraging people to trust the machine's choice over their own, as long as it isn't obviously nonsense, and the cumulative effect could be rather large after a while.

[1]: https://diff.wikimedia.org/2021/11/16/content-translation-to...

[2]: https://www.mediawiki.org/wiki/Content_translation

Re: No Language Left Behind

#113
post #106

I'll believe it when I actually see it. I'm a native of a reasonably small language spoken by about a million people and never have I ever seen a good automatic translation for it. The only translations that are good are the ones that have been manually entered, and those that match the structure of the manually entered ones. I think the sentiment is laudable and wish godspeed to the people working on this, but for t…

I speak a medium-resource language with 11 million speakers. Google Translate works so poorly with it that translations are often nonsensical. But DeepL works so well with it that translations are often indistinguishable from native speaking translations. I'm a big believer that the model can make a huge difference.

DeepL seems to handle grammar a bit better (ex. run-on sentences) but for whatever reason, it struggles with basic vocabulary sometimes. Also, when it does make mistakes, they change the meaning subtly enough to render the translation unusable.

Re: No Language Left Behind

#114
post #2

I'm not entirely sure why low resource languages are seen as such a high priority for AI research. It seems that by definition there's little payoff to solving translation for them.

Aside from the fact that being able to generalise a model with very little training data is an important AI research problem to solve, language death is a serious concern and is being accelerated due to the fact that many languages are not supported at all by modern technology (leading to "prestige language" pressures that are a known cause of historical language death).

For instance, Icelandic is not supported by any modern smartphone platform which has lead to Icelandic natives communicating with each other in English and very little information is translated to Icelandic[1,2].

That being said, I am worried that having translations that are "too good" could also act to accelerate language death as the importance of keeping languages alive will seem less significant (to non-language-nerds) if we can translate works written in that language to any other language with very small datasets. Luckily I'm not convinced that AI models will be able to produce convincing and consistent translations for a long time -- languages are so different in so many ways that I can't see how adding more dimensions and parameters to a model would account for them.

[1]: https://youtu.be/qYlmFfsyLMo?t=141 [2]: https://www.nytimes.com/2017/04/22/world/europe/iceland-icel...

Re: No Language Left Behind

#115
post #19

Jeff Wang here with my fellow Meta AI colleague Angela Fan from No Languages left Behind, seeing the comments flowing through. If you want to ask us anything, go for it!

Thank you for your exciting work and for coming onto HN to respond to questions. I am a former professional translator (Japanese to English) and am now supervising research at the University of Tokyo on the use of machine translation in second-language education. As I have written in a few papers and essays [1], advances in MT have raised serious questions for language teachers. The ready availability of MT today, in…

I've learned two languages with the help of MT. I'm sure you've interviewed people like me, but I get excited about the potential of MT for language learning, so I'd like to share my thoughts.

When I learned Spanish, I spent a lot of time chatting on Facebook with native speakers, and using Google Translate as "training wheels" to help me formulate sentences, and understand words and phrases I hadn't learned yet. It worked pretty well at the time (2012) except in cases of slang and typos that Google couldn't handle. I also used it a lot to help me translate blog posts from English to Spanish. Eventually, I graduated from the training wheels and was able to use Spanish fluently without the help of MT. More than once, while not using MT, I was told that I spoke Spanish with a "Google Translate accent", which I'm sure was more of a reference to my grammar than my accent, since my spoken practice was 100% with native speakers.

When I learned Hungarian (2019-now), at the beginning, Google Translate wasn't good enough to use for much more than getting a rough understanding of formal text, so I learned in a more traditional way at a school and with native speakers. Then the pandemic prevented me from doing both of those. I started chatting with native speakers on Facebook, but it was very difficult without MT and involved a lot of asking my conversation partners for translations and explanations. Progress was frustratingly slow. Then I discovered DeepL's MT, which was extremely good with Hungarian. I started using for chat conversations and emails, and people were shocked that I was managing to communicate with them so fluently. My progress of actually learning the language for myself accelerated dramatically. I've become conversational (B2/C1) in Hungarian in 2.5 years with very little in-person practice. Often, it takes native English speakers 5 years of in-person practice to reach that level. I'm convinced that MT played a key role in my ability to learn quickly.

When I use MT, I have a simple rule, that I have to understand each word of a translation before I send it. So I carefully read the translation, making sure that I understand each word. Sometimes that means I have to look up individual words/grammar before sending a message (I often use wiktionary for that, because it shows etymology), and other times, it means that I'll replace unfamiliar words or phrases in a translation with words and phrases from my own vocabulary. Over time I rely on MT less and less because my own vocabulary becomes stronger. I really believe that they key to learning a language quickly is to start USING the language as quickly as possible. Once you're using a language, your brain automatically starts picking up the skill. With traditional language learning, using a language can be very difficult in the beginning until you've reached a conversational level, but with MT, you can start using a language before you know everything.

For Spanish, I almost never use MT anymore. Sometimes I use it as a quick dictionary for an unfamiliar word, but my Spanish level is C2 and I use Spanish every day so it feels natural. I'm not ever translating in my head anymore.

For Hungarian, I'm still using MT often, but I don't need it during conversation (either written or spoken). Besides using it to translate things I don't know, I also find it useful for inputting Hungarian characters that are a pain to type with my US keyboard, and for conjugating words correctly when I know the root but am struggling for the correct ending. Often I'll know what I want to say in Hungarian, but I'll open DeepL and type in English, then adjust the translation to use the words I want before I copy and paste the Hungarian. I'm essentially using MT as a guide to help me craft my sentences even when I know what I want to say.

In summary, MT is awesome for language learning and for assisting language skills in development.

Re: No Language Left Behind

#116

What are hardware requirements to run this? I see the mixture model is ~ 300 GB and was trained on 256 GPUs. I assume distilled versions can easily be run on one GPU.

We release several smaller models as well: https://github.com/facebookresearch/fairseq/tree/nllb/exampl... that are 1.3B and 615M parameters. These are usable on smaller GPUs. To create these smaller models but retain good performance, we use knowledge distillation. If you're curious to learn more, we describe the process and results in Section 8.6 of our paper: https://research.facebook.com/publications/no-language-…

"All models are licensed under CC-BY-NC 4.0" :

So, to clarify, does this mean that companies cannot use these models in the course of business, or is it more about selling the translation results directly?

Re: No Language Left Behind

#117
post #19

Jeff Wang here with my fellow Meta AI colleague Angela Fan from No Languages left Behind, seeing the comments flowing through. If you want to ask us anything, go for it!

I'm curious how much work it takes to prepare training data for a language. From anecdotal experience, I've always been able to learn some basic survival skills in a new language by studying the translations of about 20 key phrases for a week or so, which give me the ability to combine them into a few hundred different phrases and survive most daily transactions. So I always imagine that training a language model is similar, just on a much larger scale. It seemed to me that there could be a standard text that includes a lot of important topics and contexts, which just needs to be manually translated into a target language and then fed to the model. I imagine it being about the size of a large book, so I imagine that adding a new language to a model would cost a similar amount to paying to have a book translated. Obviously the size of the input text would have an effect on how good the model's translations are, and domain specific translations would require more specific input. While having a full translation of an entire library seems like a good way to train a model that's used to translate everything, it seems like a small percentage of the library would be enough to produce native-level translations for most domains.

How far off are my intuitions on this? What are the costs of adding a new language to a model like this? Is there a ballpark dollar amount per language?

Re: No Language Left Behind

#118

Earlier quoted context omitted.

> But if we don't work on it, it's not gonna happen. That’s exactly right. There’s too much bias in society that if something isn’t perfect, then why bother? Nothing is perfect, so with that attitude there can be no progress. Thank you for doing important work!

Personally I'm hoping that globalisation prunes out as many languages as possible before we end up with brain implants automatically translating everything for us and no one can communicate without these chips.

Becoming bilingual is one thing. Completely extinguishing a language is a totally different matter. It is usually associated with migrating away from the geographic area of the language and/or physically losing speakers (old age, wars, genocides, etc.)

You can check the list at https://en.wikipedia.org/wiki/List_of_languages_by_time_of_e...

To make an example and be blunt: I do not expect any European country official language to get extinct anytime during our lifespan unless that country gets destroyed, which obviously won't be a good thing.

As for brain implants, I won't hold my breath.

Re: No Language Left Behind

#119

So they have a system that can translate to languages for which there isn't as much data as English, Spanish, etc. Waiting for a Twitter thread from a native speaker of one of these "low resource languages" to let us know how good the actual translations are. Cynically, I'd venture that they hired some native speakers to cherry pick their best translations for the story books. But mostly this just seems like a nice b…

If you're curious to try the system yourself, it's actually being used to help Wikipedia editors write articles for low-resource language Wikipedias: https://twitter.com/Wikimedia/status/1544699850960281601

How is the license of the models (CC NC) compatible with licenses used in Wikipedia? Did you sign an special agreement with the Wikimedia Foundation?

Re: No Language Left Behind

#120
post #106

Earlier quoted context omitted.

I speak a medium-resource language with 11 million speakers. Google Translate works so poorly with it that translations are often nonsensical. But DeepL works so well with it that translations are often indistinguishable from native speaking translations. I'm a big believer that the model can make a huge difference.

DeepL seems to handle grammar a bit better (ex. run-on sentences) but for whatever reason, it struggles with basic vocabulary sometimes. Also, when it does make mistakes, they change the meaning subtly enough to render the translation unusable.

Google Translate does the same in many languages, to the point that it will often reverse the meaning of a sentence. I honestly feel like these tools are still mostly useful when you don't really need to know what the text means.
Post reply on HN