Live data from Hacker News

A.I. Doesn't Get Black Twitter

inverse.com

91–100 of 105 posts

Re: A.I. Doesn't Get Black Twitter

#91

Earlier quoted context omitted.

> Do you think written Standard English is less ambiguous than any other English, or written language? Absolutely. It follows a strict set of formal rules and you can readily find sources to learn those rules. Heck, copyeditors regularly debate the specific nuances of how to consistently write Standard English. Other dialects (and spoken language) are far less formalized. Where is the AP Stylebook or Strunk & White f…

Can you give an example where Strunk and White or a style guide vhelp to disambiguate some expression? I can't think of any.

The Oxford comma is the most obvious case I can think of.

Re: A.I. Doesn't Get Black Twitter

#92

Earlier quoted context omitted.

Thanks for the link. I found the regional breakdowns interesting. That being said, I think this research emphasizes that AAVE is anything but standardized . That's not meant as a pejorative statement: it's just acknowledging that, like most languages in history, AAVE has not gone through a process of codification and standardization to formalize it.

Sure, but formal rules are not equivalent to a prescriptive grammar. AAVE has formal, consistent rules, described by linguists. You can make mistakes in AAVE just as in standard English (see: "African American Vernacular English Is Not Standard English with Mistakes" https://web.stanford.edu/~zwicky/aave-is-not-se-with-mistake... ) like most languages in history, AAVE has not gone through a process of codification an…

If not the process of standardization, what differentiates formal and informal rules?

Re: A.I. Doesn't Get Black Twitter

#93

Earlier quoted context omitted.

I'm looking for formal rules. I'm not saying AAVE doesn't exist, I'm saying its rules (such as they are) are not formalized. Tons of rule books exist for Standard English and part of learning it is to memorize the correct rules. I don't believe any formal equivalent exists for AAVE.

"Formal rules", in the context you've chosen to speak in, are defined by this upstream comment: > NLP is not very good with standard English yet and usually doesn't generalize from topic to topic. Dialects and other languages - especially those without formal rules - will come when we can deal with standard English. The rules you're talking about, that get printed in books and studied, are not linguistic rules. Cruci…

> Crucially, this means they are not widely observed in printed standard English, which in turn means they can't be relevant to training a language model to understand printed standard English.

I agree that they're not widely observed in written English, but they are consistently observed in the WSJ, which was the origin of this entire debate.

As lqdc13 pointed out, NLP still isn't even consistently good at understanding standard English. One could reasonably posit that that's due to the inherent ambiguity and inconsistency of most writing and that focusing on a narrower, standardized document corpus (the WSJ) you could get better initial results. What, exactly, is controversial about that? Do you really think that the language of the WSJ is no more consistent and formalized than the language of Twitter users?

Re: A.I. Doesn't Get Black Twitter

#94
post #79
post #75

Earlier quoted context omitted.

Not talking about the algorithm. It's the data that are biased. The choice of algorithm just determines how interpretable those biases actually are.

> It's the data that are biased. Which data are you referring to? In most cases, the training data isn't human-generated, and if it is, we usually want to match human behavior as close as possible.

Virtually all data used to predict crimes or recidivism is fraught with human bias, for example. Not sure that we want to reproduce the bias of criminal justice system in any prediction problem involving this type of data.

Read anything written by Solon Barocas: http://solon.barocas.org/

Re: A.I. Doesn't Get Black Twitter

#95

Earlier quoted context omitted.

Sure, but formal rules are not equivalent to a prescriptive grammar. AAVE has formal, consistent rules, described by linguists. You can make mistakes in AAVE just as in standard English (see: "African American Vernacular English Is Not Standard English with Mistakes" https://web.stanford.edu/~zwicky/aave-is-not-se-with-mistake... ) like most languages in history, AAVE has not gone through a process of codification an…

If not the process of standardization, what differentiates formal and informal rules?

Formality, if present, can be equally derived from observation and description of usage, using the linguistic analysis of the kind found in the articles I linked above. Language, when viewed in total, has shared elements and patterns that supersede the narrow focus of prescriptive impositions. It's why linguists can observe things such as that the dropped copula exists in both AAVE and Russian.

Re: A.I. Doesn't Get Black Twitter

#96
post #13

Earlier quoted context omitted.

Er, how are those couple doodles examples of the empirically opposite problem? They have nothing to do with AI. (Also... is any representation of people of color an instance of anti-white bias? Because otherwise, I don't see what the problem is with those doodles.) Also: "approximately all of Google's most valuable users (or any other ad-based platform) are non-black"... "most valuable users"? Really? Putting aside w…

The problem opposite of artificial intelligence unintentionally erasing blacks is a company's built-in intelligences intentionally erasing whites, minimizing their achievements, and deconstructing their identities. Whites are not "allowed" to have achievements; they must be couched in terms of general social achievements which of course are (according to the doodles) aspirationally majority non-white, and ideally non…

Let's not have race rants on HN, please.

Re: A.I. Doesn't Get Black Twitter

#97

Earlier quoted context omitted.

>Do you think written Standard English is less ambiguous than any other English Definitely. Things are clearer when you write formally and strictly adhere to grammatical rules. When I'm speaking casually I omit words and (ab)use punctuation much differently, and while that is perfectly acceptable to do, it is more complex and ambiguous to parse. As far as I'm aware the same pattern exists in a lot of languages and di…

To your comment that was detected as a duplicate: >I don't think there is anything "formal" about Standard English. "Standard" English as she is spoke? No. But a3n was specifically talking about the formal version often called "Standard written English". It is very prescriptive and less flexible, which makes it easier to parse. > Can you give an example of what you're talking about? Which grammatical rules do you hav…

I'm not trying to be dense, but I still don't see it. Punctuation helps with identifying boundaries of clauses, whether they're dependent or not, whether a sentence is a question. Is this what NLP struggles with? I don't think so. One of the big unsolved problems is understanding the referents of pronouns (and other proforms, like "I will [do so] too"). Style guides and generally prescriptivism does not help that.

Another ambiguity is parse ambiguities like " I saw the girl with the telescope." What does "formal" English say about that?

Re: A.I. Doesn't Get Black Twitter

#98
post #96

Earlier quoted context omitted.

The problem opposite of artificial intelligence unintentionally erasing blacks is a company's built-in intelligences intentionally erasing whites, minimizing their achievements, and deconstructing their identities. Whites are not "allowed" to have achievements; they must be couched in terms of general social achievements which of course are (according to the doodles) aspirationally majority non-white, and ideally non…

Let's not have race rants on HN, please.

How is this a "race rant"?

It's simply a well written and concise opinion (that's the opposite of a rant), completely relevant to the article in question, the contents of which, moderator 'dang' disagrees with. This is easily seen by his history of comments which exclusively censor and censure only one viewpoint.

And this is fine to hold an opinion as the expense of keeping an open mind. It is fine to use your power as a mod to advance your opinion. What is not fine is to hide your agenda underneath feigned impartiality. Your comment should read "White people can not complain about being snubbed by tech companies on this forum". It would be a far more direct and honest approach.

Re: A.I. Doesn't Get Black Twitter

#99
post #71

Earlier quoted context omitted.

In that paper it sounds like the algorithm is finding true relationships that the author doesn't like. What's the problem with the algorithm?

I don't understand what you mean by "the algorithm" or "true relationships". Do you know what I mean? There's truth as in what the data says, and there's truth as in what's underlying before social processes " corrupt " the truth. This paper examines ways to get at the latent truth. Do you think that's a useful endeavor? It's dismissive to say it has anything to do with what the "author likes".

> and there's truth as in what's underlying before social processes " corrupt " the truth.

Ah, so you're defining "truth" as "whatever I imagine the world would be like if society worked the way I wished it did". There's another word for that; "fantasy".

Re: A.I. Doesn't Get Black Twitter

#100
post #94
post #79

Earlier quoted context omitted.

> It's the data that are biased. Which data are you referring to? In most cases, the training data isn't human-generated, and if it is, we usually want to match human behavior as close as possible.

Virtually all data used to predict crimes or recidivism is fraught with human bias, for example. Not sure that we want to reproduce the bias of criminal justice system in any prediction problem involving this type of data. Read anything written by Solon Barocas: http://solon.barocas.org/

How is recidivism data biased? I'm sure that the information gleaned from parole officers and cops might be biased, but as long as the ML system is trained on whether or not someone actually reverted to committing crimes, it should be able to detect bias on the part of P.O.s and other functionaries and give a more accurate determination as to someone's chances of recidivism.
Post reply on HN