Live data from Hacker News

A.I. Doesn't Get Black Twitter

inverse.com

1–10 of 105 posts

Re: A.I. Doesn't Get Black Twitter

#3
"Approximately 8 of the 319 million people in the United States read the Wall Street Journal, about 2 percent of the population. If you look at the language — standardized English — being fed into many natural language processing units, it’s based on the language of that 2 percent. "

It's hard to take an article seriously when it opens with a logical fallacy. Yes, the WSJ uses standard English and only 8 million people read (subscribe to?) it. That does not mean that only 8 million people in the US use so-called standard English. There are also a lot of people who don't subscribe to that particular publication but still speak or write in "standard" English and that group is much larger than the group of WSJ readers.

Re: A.I. Doesn't Get Black Twitter

#4
post #3

"Approximately 8 of the 319 million people in the United States read the Wall Street Journal, about 2 percent of the population. If you look at the language — standardized English — being fed into many natural language processing units, it’s based on the language of that 2 percent. " It's hard to take an article seriously when it opens with a logical fallacy. Yes, the WSJ uses standard English and only 8 million peop…

Nah. It'd be wrong in the way you suggest if they said it's the language of only that 2%, I think.

It IS "the" language of that two million. Just not ONLY those two million.

Re: A.I. Doesn't Get Black Twitter

#5
This is an interesting ethical challenge.

If the training corpus for a machine learning model contains stereotypes and biases, then the output of the model will reflect those prejudices.

When that model is used in large-scale applications, it will not just repeat those biases, it will amplify them. It propagates biased language, which feeds back into the model in a self-reinforcing vicious cycle of increasing bias.

Should we attempt to detect prejudices in machine learning models and actively counter them, creating a self-reinforcing virtuous cycle of balance and tolerance?

----

As an aside, for an article about language, this one has a lot of typos. Doesn't anyone proofread before posting?

Re: A.I. Doesn't Get Black Twitter

#6
post #5

This is an interesting ethical challenge. If the training corpus for a machine learning model contains stereotypes and biases, then the output of the model will reflect those prejudices. When that model is used in large-scale applications, it will not just repeat those biases, it will amplify them. It propagates biased language, which feeds back into the model in a self-reinforcing vicious cycle of increasing bias. S…

Especially as it relates to this: "This means that blogs or websites that employ African-American language could actually be pushed down in search results because of Google’s language processing."

Isn't the ideal system to not assume anything about the content it's analyzing, but to base the model on behavioral patterns instead? In the case of Google ranking, either a site has traffic/low bounce/backlinks/social cred or it doesn't--why is Google attempting to 'read' and analyze the content itself? (Other than to serve ads of course)

Re: A.I. Doesn't Get Black Twitter

#7
It's interesting that they frame the problem as "woe be to the black users, the AI might discriminate against them". First, approximately all of Google's most valuable users (or any other ad-based platform) are non-black. Secondly, the problem is empirically the opposite.

https://www.google.com/search?q=white+man+white+woman&biw=13...

https://www.google.com/doodles/veterans-day-2015

http://www.cbc.ca/news/trending/google-doodle-juno-reaches-j...

Re: A.I. Doesn't Get Black Twitter

#9
post #6
post #5

This is an interesting ethical challenge. If the training corpus for a machine learning model contains stereotypes and biases, then the output of the model will reflect those prejudices. When that model is used in large-scale applications, it will not just repeat those biases, it will amplify them. It propagates biased language, which feeds back into the model in a self-reinforcing vicious cycle of increasing bias. S…

Especially as it relates to this: "This means that blogs or websites that employ African-American language could actually be pushed down in search results because of Google’s language processing." Isn't the ideal system to not assume anything about the content it's analyzing, but to base the model on behavioral patterns instead? In the case of Google ranking, either a site has traffic/low bounce/backlinks/social cred…

[deleted]

Re: A.I. Doesn't Get Black Twitter

#10
post #6
post #5

This is an interesting ethical challenge. If the training corpus for a machine learning model contains stereotypes and biases, then the output of the model will reflect those prejudices. When that model is used in large-scale applications, it will not just repeat those biases, it will amplify them. It propagates biased language, which feeds back into the model in a self-reinforcing vicious cycle of increasing bias. S…

Especially as it relates to this: "This means that blogs or websites that employ African-American language could actually be pushed down in search results because of Google’s language processing." Isn't the ideal system to not assume anything about the content it's analyzing, but to base the model on behavioral patterns instead? In the case of Google ranking, either a site has traffic/low bounce/backlinks/social cred…

> Isn't the ideal system to not assume anything about the content it's analyzing ...

Because Google no longer just searches for literal strings of text, ranked by links.

They now infer your meaning and find results that match your intent -- even if the exact search string doesn't exist in the search results.

For more on this, read about Hinton's work on "thought vectors" [1].

So if Google can't determine the meaning of African American language as well as it can standard English, then that content is likely to be less visible in search results.

[1] https://www.theguardian.com/science/2015/may/21/google-a-ste...

Post reply on HN