Live data from Hacker News

A.I. Doesn't Get Black Twitter

inverse.com

51–60 of 105 posts

Re: A.I. Doesn't Get Black Twitter

#51
post #11

Earlier quoted context omitted.

Nah. It'd be wrong in the way you suggest if they said it's the language of only that 2%, I think. It IS "the" language of that two million. Just not ONLY those two million.

Then why bring up WSJ at all? What about all the other papers? and what about all the other sources of training data for NLP research? WSJ dataset is not the only such dataset. NLP is not very good with standard English yet and usually doesn't generalize from topic to topic. Dialects and other languages - especially those without formal rules - will come when we can deal with standard English.

All language has rules - there is no language 'without formal rules'

Re: A.I. Doesn't Get Black Twitter

#53
post #32

Earlier quoted context omitted.

I get the same feeling - that Google is often missing the meaning and giving me unrelated keyword matches.

It's even worse than that - it explicitly ignores quoted keywords in many cases. The only intent they care about is what guides you towards their interests.

my favorite is from today when searching for "comparison prohibited benchmarks" google decide to give me results to "comparison ALLOWED benchmarks"

Re: A.I. Doesn't Get Black Twitter

#54
post #30
post #5

This is an interesting ethical challenge. If the training corpus for a machine learning model contains stereotypes and biases, then the output of the model will reflect those prejudices. When that model is used in large-scale applications, it will not just repeat those biases, it will amplify them. It propagates biased language, which feeds back into the model in a self-reinforcing vicious cycle of increasing bias. S…

> If the training corpus for a machine learning model contains stereotypes and biases, then the output of the model will reflect those prejudices. Depends what you mean by "the corpus". Let's take the example of loan applications. Maybe the training data contains credit scores, biographical information, and personal essays about the applicants from loan officers. Let's say loan officers tend to use more negative lang…

> If the output examples we are training on are true, the ML algorithm won't adopt any incorrect biases.

This is rarely the case when working wtih real data, and thus inspecting whether our models are biased against protected classes is probably one of the most important things an ML practitioner should do.

Re: A.I. Doesn't Get Black Twitter

#55
post #11

Earlier quoted context omitted.

Then why bring up WSJ at all? What about all the other papers? and what about all the other sources of training data for NLP research? WSJ dataset is not the only such dataset. NLP is not very good with standard English yet and usually doesn't generalize from topic to topic. Dialects and other languages - especially those without formal rules - will come when we can deal with standard English.

All language has rules - there is no language 'without formal rules'

What are the "formal rules" of AAVE?

Re: A.I. Doesn't Get Black Twitter

#56

Is this a result of the AI being written by a bunch of non-black people? It seems to me that this problem is more of a reflection of the people working on the code.

This is absolutely one of the reasons. See Google's computer vision system classifying black folks as gorillas: http://blogs.wsj.com/digits/2015/07/01/google-mistakenly-tag...

Re: A.I. Doesn't Get Black Twitter

#57
post #10
post #6

Earlier quoted context omitted.

Especially as it relates to this: "This means that blogs or websites that employ African-American language could actually be pushed down in search results because of Google’s language processing." Isn't the ideal system to not assume anything about the content it's analyzing, but to base the model on behavioral patterns instead? In the case of Google ranking, either a site has traffic/low bounce/backlinks/social cred…

> Isn't the ideal system to not assume anything about the content it's analyzing ... Because Google no longer just searches for literal strings of text, ranked by links. They now infer your meaning and find results that match your intent -- even if the exact search string doesn't exist in the search results. For more on this, read about Hinton's work on "thought vectors" [1]. So if Google can't determine the meaning…

Try some sample "AAVE" texts in Google, it does a decent job of capturing intent.

My guess is that it's unlikely to provide sources that are in one of these dialects, as the link density will be low.

Re: A.I. Doesn't Get Black Twitter

#58
post #3

"Approximately 8 of the 319 million people in the United States read the Wall Street Journal, about 2 percent of the population. If you look at the language — standardized English — being fed into many natural language processing units, it’s based on the language of that 2 percent. " It's hard to take an article seriously when it opens with a logical fallacy. Yes, the WSJ uses standard English and only 8 million peop…

Nah. It'd be wrong in the way you suggest if they said it's the language of only that 2%, I think. It IS "the" language of that two million. Just not ONLY those two million.

You can read the WSJ without going about your everyday life forming sentences the same way. It's a terrible claim no matter how you interpret it.

Re: A.I. Doesn't Get Black Twitter

#59
post #3

"Approximately 8 of the 319 million people in the United States read the Wall Street Journal, about 2 percent of the population. If you look at the language — standardized English — being fed into many natural language processing units, it’s based on the language of that 2 percent. " It's hard to take an article seriously when it opens with a logical fallacy. Yes, the WSJ uses standard English and only 8 million peop…

It's hard to take a poster seriously when they show willful ignorance of the way language serves as a marker of class, ethnicity and race, and instead have to manufacture quibbles to try to discredit an article that might cause slightly less comfort with the notion of the Glorious AI Future™ -- Coming Soon™!

I think your post is the most aggressive instance of putting strawman words in someone else's mouth I've seen all month.

Re: A.I. Doesn't Get Black Twitter

#60
post #10

Earlier quoted context omitted.

> Isn't the ideal system to not assume anything about the content it's analyzing ... Because Google no longer just searches for literal strings of text, ranked by links. They now infer your meaning and find results that match your intent -- even if the exact search string doesn't exist in the search results. For more on this, read about Hinton's work on "thought vectors" [1]. So if Google can't determine the meaning…

But if I searched in the dialect of the content I would get the content. This is simply segregating the internet. Not "pushing down" rankings. As an aside - this is a slippery slope conversation. Some people believe language requires rules. Other, more special people, think language is fluid and anarchist. The latter of those two do not make search engines.

For your aside, I don't think it's true. Search engine makers are well aware that language usage is what it is, and not what the thesaurus or grammar book says it is. Unfortunately, using machine learning to work out synonyms means you're going to miss out on minority usages, such as specialized academic jargon, or communities with millions of speakers which have slang and conversational grammar unlike "standard" English.
Post reply on HN