When I was studying Computational Linguistics I kept running into the unspoken question: given that Google Translate already exists, what is even the point of all of this? We were learning all these ideas about how to model natural language and tag parts of speech using linguistic theory so we could eventually discover that utopian solution that would let us feed two language models into a machine to make it perfectl…
Because for the other 20 percent it's plainly -not- good enough. It can't even produce an acceptable business letter in a resource-rich target language, for example. It just gets you "a good chunk of the way there."
And there's no evidence that either (1) throwing exponentially more data at the problem with see matching gains in accuracy or (2) this additional data will even be available.