> An AI model taught to view racist language as normal is obviously bad. The researchers, though, point out a couple of more subtle problems. One is that shifts in language play an important role in social change; the MeToo and Black Lives Matter movements, for example, have tried to establish a new anti-sexist and anti-racist vocabulary. An AI model trained on vast swaths of the internet won’t be attuned to the nuances of this vocabulary and won’t produce or interpret language in line with these
new cultural norms.This is a pretty superficial take on what is an extremely interesting sociological topic. (To be clear, I’m referring to the article, not the underlying paper which we don’t have.) Obviously just because social movements “have tried to establish ... vocabulary” doesn’t meant that vocabulary has become a “new cultural norm.” Plenty of such efforts end up being cultural dead-ends.
Take for example a term like “LatinX.” This term has been proposed and is used by certain people, but is extremely unfamiliar and often alienating to Latinos themselves: https://www.vox.com/2020/11/5/21548677/trump-hispanic-vote-l... (“[O]nly 3 percent of US Hispanics actually use it themselves.... The message of the term, however, is that the entire grammatical system of the Spanish language is problematic, which in any other context progressives would recognize as an alienating and insensitive message.”).
The article hand-waves away a deeply interesting question: What should an AI do here? Should AI reflect society, or be a vehicle for accelerating change? It seems at least reasonable to say that the AI should reflect what people actually say, in which case a big training dataset is appropriate, instead of what some experts decide that people should say. In some contexts, for example with “LatinX,” researchers seeking to enhance inclusivity could instead end up imposing a kind of racist elitism. (People without college educations—which disproportionately comprises immigrants and people of color—tend to be less knowledgeable about and slower to adopt these changes in vocabulary.)
The paper seems to imply that AIs should not reflect “social norms” but that training data should be selected to accentuate “attempt[ed]” shifts in such norms. Maybe that’s true, but it doesn’t seem obviously true. To return to the example above, is some Google AI generating the phrase “LatinX” (which 3/4 of Latinos have never even heard of: https://www.pewresearch.org/hispanic/2020/08/11/about-one-in...) in preference to “Latino” or “Hispanic” actually the desired result?