Live data from Hacker News

A.I. Doesn't Get Black Twitter

inverse.com

41–50 of 105 posts

Re: A.I. Doesn't Get Black Twitter

#41
post #32
post #24

Earlier quoted context omitted.

I can't say I would think that Google does a particularly good job at interpreting the intent of my search queries. I often need to reformulate them to get what I'm after. But I doubt that my searches are similar to those of the general public so that is understandable.

I get the same feeling - that Google is often missing the meaning and giving me unrelated keyword matches.

It's even worse than that - it explicitly ignores quoted keywords in many cases. The only intent they care about is what guides you towards their interests.

Re: A.I. Doesn't Get Black Twitter

#42
post #10

Earlier quoted context omitted.

> Isn't the ideal system to not assume anything about the content it's analyzing ... Because Google no longer just searches for literal strings of text, ranked by links. They now infer your meaning and find results that match your intent -- even if the exact search string doesn't exist in the search results. For more on this, read about Hinton's work on "thought vectors" [1]. So if Google can't determine the meaning…

"They now infer your meaning and find results that match your intent -- even if the exact search string doesn't exist in the search results." Bing seems to do the same thing. I really wish there were a grep for the web because both engines get it wrong even when I try to exclude results.

Try verbatim mode with google: http://paste.click/SWqasi

It's basically grep for the web, since it uses the actual query instead of the meaning of it.

Re: A.I. Doesn't Get Black Twitter

#43
For context, Olga Russakovsky is the lead author behind the ImageNet Large Scale Visual Recognition Challenge, This is one of the most popular computer vision benchmarks in the world. Olga has expertise dealing with issues of dataset bias.

A large number of ImageNet classes are dogs and birds, so most of the representational power in CNNs trained on this set are devoted to distinguishing between similar-looking dog breeds and bird species.

Re: A.I. Doesn't Get Black Twitter

#45
This, I suspect but do not know, may be a failing of the current direction of NL AI solutions. Selecting and using datasets to train NLP systems is a hard thing, especially if you truly treat it as a pattern to be learned rather than an enhancement of an underlying description of the particular language you're trying to learn.

Are we headed towards a time when deeply trained networks seem to work in an acceptably large number of cases, but how and why are unknown and susceptible to fatal flaws that rarely but catastrophically reveal themselves? I suspect that is the case with the current path of machine learning but hope I am wrong. It just seems to me that the inevitable results of certain efforts in machine learning are magic boxes that "work" but no one understands why or how, which smacks of the days of medicine prior to an understanding of the germ theory of disease where certain efforts at sanitation due to the belief in "ill humors" did improve health but were based on utterly false underlying theories, and those successes tended to reinforce other false solutions that did not improve health but in fact was harmful.

Re: A.I. Doesn't Get Black Twitter

#48
post #10
post #6

Earlier quoted context omitted.

Especially as it relates to this: "This means that blogs or websites that employ African-American language could actually be pushed down in search results because of Google’s language processing." Isn't the ideal system to not assume anything about the content it's analyzing, but to base the model on behavioral patterns instead? In the case of Google ranking, either a site has traffic/low bounce/backlinks/social cred…

> Isn't the ideal system to not assume anything about the content it's analyzing ... Because Google no longer just searches for literal strings of text, ranked by links. They now infer your meaning and find results that match your intent -- even if the exact search string doesn't exist in the search results. For more on this, read about Hinton's work on "thought vectors" [1]. So if Google can't determine the meaning…

But if I searched in the dialect of the content I would get the content. This is simply segregating the internet. Not "pushing down" rankings.

As an aside - this is a slippery slope conversation. Some people believe language requires rules. Other, more special people, think language is fluid and anarchist. The latter of those two do not make search engines.

Re: A.I. Doesn't Get Black Twitter

#49
post #42

Earlier quoted context omitted.

"They now infer your meaning and find results that match your intent -- even if the exact search string doesn't exist in the search results." Bing seems to do the same thing. I really wish there were a grep for the web because both engines get it wrong even when I try to exclude results.

Try verbatim mode with google: http://paste.click/SWqasi It's basically grep for the web, since it uses the actual query instead of the meaning of it.

Thanks, I did not know that was an option. It should make a few different queries quite a lot nicer. I was starting to get scared that I was wishing for the old Alta Vista.

Re: A.I. Doesn't Get Black Twitter

#50
post #3

"Approximately 8 of the 319 million people in the United States read the Wall Street Journal, about 2 percent of the population. If you look at the language — standardized English — being fed into many natural language processing units, it’s based on the language of that 2 percent. " It's hard to take an article seriously when it opens with a logical fallacy. Yes, the WSJ uses standard English and only 8 million peop…

It's hard to take a poster seriously when they show willful ignorance of the way language serves as a marker of class, ethnicity and race, and instead have to manufacture quibbles to try to discredit an article that might cause slightly less comfort with the notion of the Glorious AI Future™ -- Coming Soon™!
Post reply on HN