Live data from Hacker News

A.I. Doesn't Get Black Twitter

inverse.com

11–20 of 105 posts

Re: A.I. Doesn't Get Black Twitter

#11
post #3

"Approximately 8 of the 319 million people in the United States read the Wall Street Journal, about 2 percent of the population. If you look at the language — standardized English — being fed into many natural language processing units, it’s based on the language of that 2 percent. " It's hard to take an article seriously when it opens with a logical fallacy. Yes, the WSJ uses standard English and only 8 million peop…

Nah. It'd be wrong in the way you suggest if they said it's the language of only that 2%, I think. It IS "the" language of that two million. Just not ONLY those two million.

Then why bring up WSJ at all? What about all the other papers? and what about all the other sources of training data for NLP research?

WSJ dataset is not the only such dataset. NLP is not very good with standard English yet and usually doesn't generalize from topic to topic. Dialects and other languages - especially those without formal rules - will come when we can deal with standard English.

Re: A.I. Doesn't Get Black Twitter

#12
post #5

This is an interesting ethical challenge. If the training corpus for a machine learning model contains stereotypes and biases, then the output of the model will reflect those prejudices. When that model is used in large-scale applications, it will not just repeat those biases, it will amplify them. It propagates biased language, which feeds back into the model in a self-reinforcing vicious cycle of increasing bias. S…

> When that model is used in large-scale applications, it will not just repeat those biases, it will amplify them. It propagates biased language, which feeds back into the model in a self-reinforcing vicious cycle of increasing bias.

Which is precisely why social media (e.g. your newsfeed) is such an awful source for news, especially when you consider how low the standards are on digital publishing these days...

Re: A.I. Doesn't Get Black Twitter

#13
post #7

It's interesting that they frame the problem as "woe be to the black users, the AI might discriminate against them". First, approximately all of Google's most valuable users (or any other ad-based platform) are non-black. Secondly, the problem is empirically the opposite. https://www.google.com/search?q=white+man+white+woman&biw=13... https://www.google.com/doodles/veterans-day-2015 http://www.cbc.ca/news/trending/go…

Er, how are those couple doodles examples of the empirically opposite problem? They have nothing to do with AI. (Also... is any representation of people of color an instance of anti-white bias? Because otherwise, I don't see what the problem is with those doodles.)

Also: "approximately all of Google's most valuable users (or any other ad-based platform) are non-black"... "most valuable users"? Really? Putting aside whether that's even an accurate assessment in the financial terms in which it was presumably intended: Do ethics stop at treating well only the people with most profit potential? Is impact on business our only arbiter?

Re: A.I. Doesn't Get Black Twitter

#14
post #7

It's interesting that they frame the problem as "woe be to the black users, the AI might discriminate against them". First, approximately all of Google's most valuable users (or any other ad-based platform) are non-black. Secondly, the problem is empirically the opposite. https://www.google.com/search?q=white+man+white+woman&biw=13... https://www.google.com/doodles/veterans-day-2015 http://www.cbc.ca/news/trending/go…

Google, probably more than any other company in the world, gets that their "most valuable users" aren't large coherent groups of any sort, Google wins by catering to everybody at once (aka "the long tail", to use a 2000s-ism). This is what drives their relentless focus on machine learning everything, because machine learning allows you to serve microscopic constituencies. You don't need machine learning to sell beer to white, suburban males.

Re: A.I. Doesn't Get Black Twitter

#16

Is this a result of the AI being written by a bunch of non-black people? It seems to me that this problem is more of a reflection of the people working on the code.

There may be some of that, but I think it has more to do with AI still being young, and it's easier to work with less ambiguous text for learning. That and everyday speech move much faster than formal writing. Now that my son has gone off to college and is less accessible as a resource, I expect to almost totally lose touch with the speech of "these kids today."

Re: A.I. Doesn't Get Black Twitter

#17
post #5

This is an interesting ethical challenge. If the training corpus for a machine learning model contains stereotypes and biases, then the output of the model will reflect those prejudices. When that model is used in large-scale applications, it will not just repeat those biases, it will amplify them. It propagates biased language, which feeds back into the model in a self-reinforcing vicious cycle of increasing bias. S…

> Doesn't anyone proofread before posting?

No.

Re: A.I. Doesn't Get Black Twitter

#18
There is nothing about a computer algorithm that renders it immune to the power structures around it. The people who train these systems don't even realize they are creating a biased model. Borrowing from postcolonial theory, this is related to the concept of epistemic violence. The way we organize human information, the way we categorize and evaluate viewpoints and ways of understanding the world is a battleground for imperialism. Intent does not matter. The algorithm has no intent, the researchers don't intend to create a vehicle of oppression, and yet it happens anyway. They have created a difference in economic utility between standard, hegemonic English, and marginalized dialects of English which inevitably will have social ramifications. If this study weren't conducted, would we have ever found out?

Re: A.I. Doesn't Get Black Twitter

#19
post #5

This is an interesting ethical challenge. If the training corpus for a machine learning model contains stereotypes and biases, then the output of the model will reflect those prejudices. When that model is used in large-scale applications, it will not just repeat those biases, it will amplify them. It propagates biased language, which feeds back into the model in a self-reinforcing vicious cycle of increasing bias. S…

So humans do this too naturally - this is how new languages develop over hundreds of years and groups evolve and split. Aside from that, I'm a little nervous about seeing machines accelerate this process.

Re: A.I. Doesn't Get Black Twitter

#20
post #8

I bet it doesn't get Scottish Twitter either. I saw this the other day and thought it was pretty funny: https://mobile.twitter.com/MarkHamiIl/status/778141129564905...

When I first moved to London, one of my co-worker was a very exuberant young Scottish woman. Her accent was near impossible for me, and she fully realised this, but at one point we came to the mutual understanding that I did in fact pick up most of what she was saying, even though it was only because she talked so much, that she didn't really need to intentionally repeat what she was saying.
Post reply on HN