Live data from Hacker News

A brief history and future of credit scores

economist.com

61–70 of 72 posts

Re: A brief history and future of credit scores

#61
post #54

Earlier quoted context omitted.

Isn't that a bit like asking how many pedestrians you can hit and still keep your license? The right answer is to try for zero. The dilemma you're describing isn't due to the law itself, but rather because the difficulty of writing law results in only the absolute worst abuses being criminalized. The abstract safe answer is to not engage in group-based discrimination at all , regardless of it seeming quite lucrative…

> Isn't that a bit like asking how many pedestrians you can hit and still keep your license? This is nothing like that. Hitting or not hitting a pedestrian is binary, and there is a clear way to determine fault. We often call car collisions "accidents", but they aren't. Someone did something wrong to cause a collision. The issue here is that you can be discriminatory "by accident". You can be 100% race blind and have…

> there is a clear way to determine fault

Yes, but criminal intent is less clear. Someone can be declared at fault for hitting a pedestrian, yet not have been criminally negligent. And even if this happens a few times from really bad luck, it's reasonable that they'd get to keep their license [0].

> You can be 100% race blind and have the best intentions in the world ... bias due to the nature of the data

The question is what data? Are you feeding things in that have a clear causal relationship with your desired result? Or are you inputting everything you can hoping to discover correlations ? The latter is essentially trying to suss out informal groups, and engaging in any group-based discrimination means you cannot claim to have "the best intentions" - regardless of whether the group can be named as a legally protected one or not. If say you're try to base mortgage underwriting on credit card purchase data, it's wholly disingenuous to claim that the resulting fallout is "accidental".

[0] In a society where cars are de facto mandatory to get around. I will disclaim that this casual attitude is a large part of what keeps roads so hostile for everyone else, but it currently is what it is. Also historically we haven't had such a likely hidden factor as phone use.

Re: A brief history and future of credit scores

#62

Earlier quoted context omitted.

> In general, the standard is "disparate impact" That standard is unworkable. What happens when an idealistic charity goes into a homeless shelter in a black neighborhood with a high rate of drug use and helps them all fill out mortgage applications, so that 80% of the lender's white applicants are married with stable middle class jobs and 80% of the black applicants are single, homeless, unemployed and suffering fro…

There is a legal defense, where you can argue that the disparate impact is caused by practical business need instead of by implicit bias. An example (using gender instead of race) is that women have less upper body strength than men, on average, so they are far less likely to meet job requirements that require being able to carry 100 pounds of equipment on their backs. Even then, however, you still have to demonstrat…

Which is why that standard is unworkable.

If you're giving loans and justifying a factor is the same as showing that it correlates with repayment rate then it's that easy to show a justification, but it will also basically always produce a "disparate impact" because repayment rate itself correlates with race, so any reasonably accurate measure of repayment rate will do the same.

A "disparate impact" standard is useless because it's routinely met even when discrimination is not actually occurring. In fact, finding no disparate impact would be highly indicative that someone was impermissibly taking race into consideration in order to fudge the numbers.

So either you have a disparate impact because that's what naturally happens with accurate predictions, or you do the expressly prohibited thing and take race directly into consideration in order to make it go away.

Imagine a human intervening to remove a factor from consideration because considering it benefits black applicants relative to other applicants, even though considering it improves accuracy overall. Would they not rightfully get their butts sued off? But what happens if it's the same thing, only this time it benefits hispanic or korean or greek applicants relative to black applicants? The only remaining option is to consider all the information you have available, which is what people are inclined to do to begin with.

Re: A brief history and future of credit scores

#63

Earlier quoted context omitted.

There is a legal defense, where you can argue that the disparate impact is caused by practical business need instead of by implicit bias. An example (using gender instead of race) is that women have less upper body strength than men, on average, so they are far less likely to meet job requirements that require being able to carry 100 pounds of equipment on their backs. Even then, however, you still have to demonstrat…

Which is why that standard is unworkable. If you're giving loans and justifying a factor is the same as showing that it correlates with repayment rate then it's that easy to show a justification, but it will also basically always produce a "disparate impact" because repayment rate itself correlates with race, so any reasonably accurate measure of repayment rate will do the same. A "disparate impact" standard is usele…

I'm not sure what you're trying to argue here. Are you saying that disparate impact liability is fraught? Everyone agrees that it is. Cases based on it hinge on whether the impact results from "legitimate business need" with no reasonable alternatives, a standard that expressly allows e.g. credit scores with strong racial or gender correlations.

So if that's what you're concerned about, the response is simple: it's not enough to simply show a correlation; in enforcing disparate impact claims (under ECOA or FHA or Title VII), regulators have to show not just the correlation, but also the illegitimacy of the (facially neutral) action, or at last that some other (facially neutral) business practice would accomplish the same goals without producing the impact.

Since this is HN, I can't discard the idea that maybe you're instead arguing that disparate impact isn't in fact the standard in US law, in which case: no, a simple Google search for "disparate impact" and any of the laws I cited in that last paragraph will quickly disabuse you of that.

Re: A brief history and future of credit scores

#64
post #46

Earlier quoted context omitted.

I just missed the deadline to edit my post, so I am replying to myself. Looking at the parent comment again, I seemed to have just restated it without adding anything new of my own. I meant to add that my reasoning for why Amazon pulled development of this tool was not just because the tool’s bias, but also because that the existence of the tool and its associated training data could open up Amazon to litigation clai…

It's interesting that nobody even bothered to check whether the bias was illicit. They found something that sounds bad and the immediate response is "OMG bad PR, pull emergency shutdown." They just assume that "women's" is coding for female candidates and not something more specific, like gender-segregated activities that may legitimately produce lower quality candidates than the equivalent integrated activities that…

To be fair, it would probably take more time and money irrespective of litigation to be sure that there was illicit bias than to just delete everything and call the exercise a failure.

Re: A brief history and future of credit scores

#65
post #47

Earlier quoted context omitted.

Thanks, economist.com is terrible

Meta comment but why do you say this yet read their content?

Their content is good but their site design is bad. Good thing we have nice users like GP who'll post the content in a readable form.

Re: A brief history and future of credit scores

#66

Earlier quoted context omitted.

> The key here being that non-financial data isn't actually useful in predicting ability to repay loans. Sure it is. Even religion correlates with loan risk. (Could be a proxy for social status and people from your tribe helping you out when you can't pay back). I could probably get a predictive model better than random guessing by mining your HackerNews comments or Facebook likes. > I'm sure it'd be used somehow by…

Cool, I look forward to your evidence and data that you're using to contradict the story in the submission, because the very article states that if you tried to pitch a loan approval algorithm with any of the things you suggested in a modern bank, you'd be laughed out of the firm.

The evidence is in the article. It tells of a US company using 10.000 data points in a jurisdiction where this is allowed.

The article cites Turner, who a few years back wrote this:

> The non-financial data was found predictive in all three outcomes examined when no other ‘traditional’ credit information was used, strongly suggesting that alternative data would be useful to lenders in underwriting the so-called ‘no-file’ or ‘no-score’ consumer who have little or no payment/credit information available.

I think his quote about "laughing out" is taken out of context. No predictive modeler will throw away informative features, because she can not distinguish noise from signal after 26 variables. That's 10 to 1% of a modern credit scoring model. It may be the perspective of a regulator though (they start drowning in noise after reviewing 100+ variables).

Yes, all data that is legal to use and predictive, will get used, if not by you, then by your competitor.

And informative variables that can not be used in the decision to give a loan, are used internally to predict if the loan will be paid back. There is more to credit scoring than the initial yes-no.

Re: A brief history and future of credit scores

#67
post #63

Earlier quoted context omitted.

Which is why that standard is unworkable. If you're giving loans and justifying a factor is the same as showing that it correlates with repayment rate then it's that easy to show a justification, but it will also basically always produce a "disparate impact" because repayment rate itself correlates with race, so any reasonably accurate measure of repayment rate will do the same. A "disparate impact" standard is usele…

I'm not sure what you're trying to argue here. Are you saying that disparate impact liability is fraught? Everyone agrees that it is. Cases based on it hinge on whether the impact results from "legitimate business need" with no reasonable alternatives, a standard that expressly allows e.g. credit scores with strong racial or gender correlations. So if that's what you're concerned about, the response is simple: it's n…

> it's not enough to simply show a correlation; in enforcing disparate impact claims (under ECOA or FHA or Title VII), regulators have to show not just the correlation, but also the illegitimacy of the (facially neutral) action, or at last that some other (facially neutral) business practice would accomplish the same goals without producing the impact.

But that's the problem.

Suppose you're evaluating whether to consider zip code in loan approvals, and that doing so improves the overall prediction rate, helps hispanic applicants, but hurts black applicants.

If you choose to stop considering zip code, you've got a disparate impact against hispanics but no business justification for doing it that way, meanwhile they can show an alternative (i.e. taking zip code into account) that serves your business goals better and reduces the disparate impact against hispanics, so not considering it may get you into trouble.

But we also have people suggesting that you shouldn't consider it because it increases the disparate impact against black applicants, and you might "have a hard time" demonstrating that business justification in court. Which is perhaps not ridiculous, because the business justification is there, but it's also in complicated algorithms that are not easy to explain to a layman.

So you have two basically reasonable alternative courses of action, either of which could arguably result in liability. This is the hallmark of unworkable legislation.

Re: A brief history and future of credit scores

#68
post #63

Earlier quoted context omitted.

I'm not sure what you're trying to argue here. Are you saying that disparate impact liability is fraught? Everyone agrees that it is. Cases based on it hinge on whether the impact results from "legitimate business need" with no reasonable alternatives, a standard that expressly allows e.g. credit scores with strong racial or gender correlations. So if that's what you're concerned about, the response is simple: it's n…

> it's not enough to simply show a correlation; in enforcing disparate impact claims (under ECOA or FHA or Title VII), regulators have to show not just the correlation, but also the illegitimacy of the (facially neutral) action, or at last that some other (facially neutral) business practice would accomplish the same goals without producing the impact. But that's the problem. Suppose you're evaluating whether to cons…

> Suppose you're evaluating whether to consider zip code in loan approvals, and that doing so improves the overall prediction rate, helps hispanic applicants, but hurts black applicants.

That's not disparate impact. The legal code does not follow exact predicate logic, where if you meet conditions A, B, and C, you violate the law. It tends to follows rules of fuzzy logic instead--that's why you'll see legal opinions that involve words like "tends to", "probably", "factors" a lot. Particularly where a strict interpretation would lead you to an apparent contradictions, the court system instead tries to find a reasonable course of action. Indeed, often merely showing that you are making a good-faith effort to comply with all applicable laws and regulations is sufficient to absolve you of penalties for failure to comply.

Ultimately, the arbiter of reasonableness isn't a blackbox oracle. It's a panel of 12 members of the general public, or perhaps a panel of 3-9 judges.

Re: A brief history and future of credit scores

#69

Earlier quoted context omitted.

Cool, I look forward to your evidence and data that you're using to contradict the story in the submission, because the very article states that if you tried to pitch a loan approval algorithm with any of the things you suggested in a modern bank, you'd be laughed out of the firm.

The evidence is in the article. It tells of a US company using 10.000 data points in a jurisdiction where this is allowed. The article cites Turner, who a few years back wrote this: > The non-financial data was found predictive in all three outcomes examined when no other ‘traditional’ credit information was used, strongly suggesting that alternative data would be useful to lenders in underwriting the so-called ‘no-f…

The US company doesn't operate in the US, and the quote wasn't taken out of context. The context is that 10,000 was way over the 10 or so maximum number of actually valuable points of data.

The whole article is about how utterly useless the vast majority of "data" ends up being, and how they are not used internally to predict if the loan will be paid back.

So you are 100% in disagreement with this article, and forgive me if I trust a nationally published periodical such as Newsweek of a throwaway commenter on the Internet.

Re: A brief history and future of credit scores

#70
post #6

Earlier quoted context omitted.

I get confused by those they claim some of these ML or NN algos are black boxes and so could be breaking the law. If you’re not inputting illegal info (like race, sex, national origin, religion, name, etc. or corollaries), then it’s not making its risk assessment on that basis. All you have to do is look at the inputs. It isn’t unknowable whether or not it’s breaking the law.

Let me try to clear your confusion: an input my seem innocent (e.g. zip code), but a zip code is likely to correlate to ethnicity and race in some regions. So even if the inputs seem legal, an ML model that’s sophisticated enough can derive illegal results that discriminate against certain populations.

I mentioned corollaries... there’s no reason zip code or lat long should be an input.
Post reply on HN