Live data from Hacker News

Semantics derived automatically from language corpora contain human-like biases

scim.ag

71–80 of 92 posts

Re: Semantics derived automatically from language corpora contain human-like biases

#71
post #69
post #58

Earlier quoted context omitted.

> put it through an algebraic transformation that subtracts known biases > Do you consider that algebraic transformation enough of a "morality system"? I would consider it a sort of morality, yes. But keep in mind that the list of "known biases" would itself be biased toward a particular goal, be it political correctness or something else.

Yes, every step of machine learning has potential bias, we know that, that's what this whole discussion is about. Nobody would responsibly claim that they have solved bias. But they should be able to do something about it without their progress being denied by facile moral relativism. If we can't agree that one can improve a system that automatically thinks "terrorist" when it sees the word "Arab" by making it not do…

[deleted]

Re: Semantics derived automatically from language corpora contain human-like biases

#72
post #40

Earlier quoted context omitted.

So, "plants are pleasant" or "insects are unpleasant" are universal truths?

The example the parent poster described is "doctors are 66% male", which could very well be true.

My point is that neither "insects are unpleasant" nor "plants are pleasant" nor "doctors are 66% male" are immutable features of the universe. They are merely snapshots of the human view of world conditions, as the world is now. "True now", but not "true forever and always".

The paper seems to advocate for designing ML systems that learn that what is "true now" may not be "true forever and always". It seems to be quite the opposite of "there are certain truths that ML systems should not learn."

Re: Semantics derived automatically from language corpora contain human-like biases

#73

Earlier quoted context omitted.

Why would we want to reproduce existing structures of oppression in mechanical form? Have you noticed how automation often vastly amplifies things? It's a short step from saying 'this model accurately reflects the bias in society' to 'that's how things are, the computer says women aren't cut out to be doctors.' Surely you are aware that in real world world people rationalize decisions they don't actually understand a…

> Why would we want to reproduce existing structures of oppression in mechanical form? If (for example) 66% of Doctors are male and 34% female then it's not reproducing "existing structures of oppression" it's inferring something about reality .

It's inferring something about reality, but what?

Suppose, for example, that I gave this same statistic to someone and then asked them to select from a pool of 100 applicants for 50 available places in medical school. Let's assume that there's an equal # of male and female applicants and that their exam results are all similar. Do you think that knowing about this 66-34 split might influence the gender balance of the final selection?

Re: Semantics derived automatically from language corpora contain human-like biases

#74
post #36

Earlier quoted context omitted.

I think you are making a distinction without a difference. If the word vectors pick up biases from wikipedia text, than for all practical purposes, they are (indirectly) absorbing stereotypes from humans. This is an expected result, but not necessarily desirable in the end.

The GP is saying that the bias isn't an attribute of the Wikipedia text, but of reality. If the reality is that only 34% of doctors are female, why is it not desirable for the machine to learn that?

Suppose that the system "learned" that marriage consisted of one man and one (or very rarely more) woman. Would that be "reality"? (In fact, I'd rather bet that it did, and I congratulate the authors on the wisdom of not advertising that fact.)

Various behavioral accidents can easily become embedded in culture, laws, and, yes, programs, at which point it stops mattering if they represent reality or "reality"; the real world will happily follow the cultural construction.

Re: Semantics derived automatically from language corpora contain human-like biases

#75

Earlier quoted context omitted.

An exercise: Words relating to insects will occur in news articles about environmentalism, crop production, rituals of rebirth, etc. Words relating to plants might occur in articles about crop destruction, the international drug trade, people getting poisoned, etc. rmxt questioned the universality of sentiment analysis. Responding by noting specific contexts, free from a clear coherent general structure, is an assert…

But it is a universal truth that humans generally find plants pleasant and insects unpleasant. And the word "pleasant" is entirely based on human preferences after all.

"...it is a universal truth that humans generally find plants pleasant..."

Ah, but those exceptions are really unpleasant.

Re: Semantics derived automatically from language corpora contain human-like biases

#76
post #66

Earlier quoted context omitted.

A black box neural network attempting to draw inferences from a human-biased dataset - potentially even more biased because it can't understand subtexts - and then verifying that conclusion through an ad-hoc set of "morality checks" entirely independent from how it reached the conclusion sounds like a recipe for disaster. That's even before the marketing people get involved and start claiming the system is free from…

I'm not sure what your objection is with regards to the independence aspect. Why would having the morality checks integrated into the "learning about the world" part be better? If you had an unwavering moral code which dictated that men and women should be treated equally, for example, why would it matter which facts are presented to you, in what order, or how you process them? Your morality would always prevent you…

Frankly, I'm not sure the "men and women should be treated equally" instruction is even possible if the data isn't processed in a way which specifically controls for the effects of gender (some of which may not be discernible from the raw inputs).

Sure, it's theoretically possible that an algorithm parsing text about medics' credentials that (e.g) positively weights male names and references to all-boys' schools and negatively weights female names and references to Girl Guides will be on average fair after an ad hoc re-ranking of all its candidates to take into account the instruction to treat male and female candidates equally. It's just unlikely to achieve this without completely reorganizing its underlying predictive model

[1]there's an interesting parallel to ongoing human arguments about how a machine should follow its "morality checks" should do this: does it ensure the subjects are "treated equally" in terms of achieving 50/50 gender ratio irrespective of the candidate pool (thus potentially skewing it massively in favour of the side with the weaker applicants), does it try to weight results so gender balance reflects historic norms (thus permanently entrenching the minority)? Or does it try to be "gender blind" by testing all its inputs for whether they're gender biased and normalising for or discarding those which are, which is basically learning everything again from scratch...

Re: Semantics derived automatically from language corpora contain human-like biases

#77

Earlier quoted context omitted.

> Why would we want to reproduce existing structures of oppression in mechanical form? If (for example) 66% of Doctors are male and 34% female then it's not reproducing "existing structures of oppression" it's inferring something about reality .

In an environment in which Blue people are banned from becoming doctors, its also inferring something about reality to conclude that 0% of Doctors are Blue. It would be entirely wrong, however, to use these inputs to infer anything whatsoever about the respective propensity of Blue and Green people to become doctors in an environment in which such a rule or idea of a rule had never existed. Obviously "structures of o…

> It would be entirely wrong, however, to use these inputs to infer anything whatsoever about the respective propensity of Blue and Green people to become doctors in an environment in which such a rule or idea of a rule had never existed.

That's fine but it isn't the goal of these algorithms. It isn't the reality that is useful for them to learn. It's a different problem to try to build some kind of "unbiased" ontology rather than just to learn about words. Feel free to research or create solutions to this other problem, it sounds interesting.

Re: Semantics derived automatically from language corpora contain human-like biases

#78

Earlier quoted context omitted.

Why would we want to reproduce existing structures of oppression in mechanical form? Have you noticed how automation often vastly amplifies things? It's a short step from saying 'this model accurately reflects the bias in society' to 'that's how things are, the computer says women aren't cut out to be doctors.' Surely you are aware that in real world world people rationalize decisions they don't actually understand a…

>structures of oppression What oppression? How are word vectors oppressing anyone? What a ridiculous claim. >Have you noticed how automation often vastly amplifies things? No, not at all. I've heard this claim on similar discussions. But I've yet to see a convincing example. Particularly with word2vec. I find it very implausible that word vectors will somehow discriminate against female doctors or whatever. >It's a s…

No one is ever going to use word vectors to figure out what genders are capable of what jobs.

Directly, no. Nobody is going to go 'ah, word2vec - a new tool with which to perpetuate patriarchal capitalism, mwuhahaha'...probably. People are weird that way.

But indirectly they certainly will. How about NPC character generators in MMORPGs? Or chatbots on social networks? Stock characters in auto-generated romance novels? The possibilities are endless.

No doubt you will these examples are ridiculous, because you seem like a rigorous scientifically minded person who would be careful not to use data in inappropriate contexts, and who would try to discount cultural or emotional factors in making strategic decisions. But you are only as good at this as your own self-awareness and willingness to acknowledge the existence of implicit bias.

And many people are quite different from you and more easily or willingly allow their judgment to be shaped by representational stereotypes. Marketing people aim to confirm their audience's worldview very closely so that consumers will be willing to identify with the commercial prompt when it arrives. Politicians and yellow journalists routinely abuse statistics to grab people's attention. And so on.

I urge you to think more about this, and in more imaginative fashion. People are often surprised by the unexpected applications of technology employed by others.

Re: Semantics derived automatically from language corpora contain human-like biases

#79

Earlier quoted context omitted.

> Why would we want to reproduce existing structures of oppression in mechanical form? If (for example) 66% of Doctors are male and 34% female then it's not reproducing "existing structures of oppression" it's inferring something about reality .

It's inferring something about reality, but what? Suppose, for example, that I gave this same statistic to someone and then asked them to select from a pool of 100 applicants for 50 available places in medical school. Let's assume that there's an equal # of male and female applicants and that their exam results are all similar. Do you think that knowing about this 66-34 split might influence the gender balance of the…

Knowing about the gender balance wouldn't influence the final selection if you programmed the selection criteria not to be influenced by the gender balance.

The whole point of training and using machines is to make more accurate, more useful decisions in a complex world.

That can't happen if we give them data that isn't borne out by reality, or tell them to ignore data that is.

Re: Semantics derived automatically from language corpora contain human-like biases

#80
This was a neat study, but the authors really do throw around a number of speculative and unsupported claims in the discussion about linguistic relativity/determinism.

> Our findings are also sure to contribute to the debate concerning the Sapir-Whorf hypothesis (17), because our work suggests that behavior can be driven by cultural history embedded in a term’s historic use. Such histories can evidently vary between languages.

No, not really. Linguistic relativity makes a claim about the direction of causation--from language to thought. This study does nothing to test the direction of causation, which is usually considered possible only with controlled experiments. This chicken and egg debate has been going on for the better part of the last century--it is not an easy problem.

> Our results also suggest a null hypothesis for explaining origins of prejudicial behavior in humans, namely, the implicit transmission of ingroup/outgroup identity information through language. That is, before providing an explicit or institutional explanation for why individuals make prejudiced decisions, one must show that it was not a simple outcome of unthinking reproduction of statistical regularities absorbed with language.

Once again, this was not an experiment capable of producing causal evidence. The authors have shown that human biases can be replicated by statistical learning of language corpora--admirable work, but nothing new here for Sapir-Whorf.

Post reply on HN