Earlier quoted context omitted.
I get the point, but why didn't you just exclude intern resumes from the training data? Do you still suspect a skewed result?
>I get the point, but why didn't you just exclude intern resumes from the training data? That was the logical next step and we started on that, but it required exporting more historic data out of the HR system and filtering out anyone who started as an intern as well. Sounds simple, but in practice it's anything but. Just for the reference, data extraction, cleaning and filtering in that project took at least an orde…
Notes on AI Bias
111–120 of 126 posts
Re: Notes on AI Bias
#112Earlier quoted context omitted.
>I get the point, but why didn't you just exclude intern resumes from the training data? That was the logical next step and we started on that, but it required exporting more historic data out of the HR system and filtering out anyone who started as an intern as well. Sounds simple, but in practice it's anything but. Just for the reference, data extraction, cleaning and filtering in that project took at least an orde…
@gambler: thank you for reading my reporting. I would love to chat confidentially to understand your perspective better. Please see my HN profile.
Gender bias was not the only issue. Problems with the data that underpinned the models’ judgments meant that unqualified candidates were often recommended for all manner of jobs, the people said. With the technology returning results almost at random, Amazon shut down the project, they said.
Re: Notes on AI Bias
#113Earlier quoted context omitted.
> Except that's exactly what it is. Using the term 'bias' has certain political motivations behind it. It's not about the term being technically untrue as it is about the term being non-neutral. For instance, here are some definitions of 'bias' I just grabbed from American Heritage: "A preference or an inclination, especially one that inhibits impartial judgment." "An unfair act or policy stemming from prejudice." "A…
Fine: it's easy to accidentally train ML models so that they will make systematic errors. Often these errors stem from systematic biases in our society, model creators should therefore be aware of the potential biases[1] that their models could reflect, and how to prevent them. [1]: With the political motivation.
No, this does also not match.
One of the easiest way to get a ML model that creates systematic errors is spam filters. If I take my spam folder with no consideration, what the filter will learn is that any language which isn't my own are spam, and that servers located outside my nation are spammers. This resembles prejudice.
The cause of this systematic error is that individual email addresses do not get ham emails uniformly from every nation and every language. Proximity warps the data. I would need to normalize the data based on language and nation if I wanted to remove those errors in the filter. Looking at it from a political perspective does not make the filter perform better, and fixing it from that side has a high risk of causing even more errors in the model.
Re: Notes on AI Bias
#114Earlier quoted context omitted.
I’m not sure what that even means if we know we can bias outcomes. Pretending there is some kind of natural state that is for the sake of being natural preferred seems odd given humans propensity to change the world to suit. I also suspect for many that ‘reality’ is really just a dog whistle for their preferred biases. Not to mention the entire issue with deriving and ought from an is.
Suppose you train an AI to predict how good people are at weight lifting, trained from a bunch of seemingly unrelated data (maybe you want to hire bouncers or construction workers). You will find that the model predicts better performance for males. You notice this, identify that men are more likely to go to the gym than wimen, and modify your data to compensate for this. But when you rerun the model men still show b…
Your post is great for the assumptions it encodes. Like what does it mean to be good at weight lifting? And that for some reason being good at weight lifting is a good proxy for being a good bouncer or construction worker?
For an off the cuff example it’s a great way to demonstrate the sort of bias we can naively introduce then defend because it’s just ‘reality’. When really it’s much more complex than identifying a relevant trait and assuming everything else falls out of it.
Re: Notes on AI Bias
#115Earlier quoted context omitted.
>Why? Pointing out a specific and concrete harm badly designed ML models cause is irresponsible? In my opinion, yes, if it leads most readers to misjudge some fundamental properties of the problem as a whole. Again, I'm not saying this article is guilty, but most are.
> In my opinion, yes, if it leads most readers to misjudge some fundamental properties of the problem as a whole. Which problem? The general statement of this problem is "models, trained on [somehow] misrepresentative data [or even technically representative data] can draw unintended conclusions that lead to harm". Specifically in this case, the harm was "the model was basically just trained to ignore all women appli…
Throwing AI at answering an ill-formed question or optimizing a process that shouldn't happen in the first place is not something that can be corrected by getting better training data.
Moreover, automation can have consequences that aren't detectable by analyzing some test set.
Re: Notes on AI Bias
#116Earlier quoted context omitted.
Suppose you train an AI to predict how good people are at weight lifting, trained from a bunch of seemingly unrelated data (maybe you want to hire bouncers or construction workers). You will find that the model predicts better performance for males. You notice this, identify that men are more likely to go to the gym than wimen, and modify your data to compensate for this. But when you rerun the model men still show b…
What I was getting at is that our important choices are about outcomes and those have nothing really to do with assumptions about reality. For example all should be equal before the law. A statement that is supposed to be true but very obviously isn’t. Your post is great for the assumptions it encodes. Like what does it mean to be good at weight lifting? And that for some reason being good at weight lifting is a good…
being a good weight lifter, means you can lift heavier weights then a less-good weight lifter. Whether this is a proxy for anything isn't relevant, because it's a purely contrived example. There are clearly jobs where physical strength (among other things) is important, and given the context of this example, there is no guarantee that a more complicated model evens out the differences.
The point of the example is, basically "there are some things which might discriminate strongly on the basis on physical traits, which might end up correlating with race/sex etc" - ask for a better model by all means, but there is no guarantee the perfect model will never correlate strongly with some political demographic, and hence be controversial.
Re: Notes on AI Bias
#117Earlier quoted context omitted.
One situation I could see leading to this result (Amazon cancelling their resume filtering software with the excuse that it 'skewed male') is that 1. The AI system accurately predicted employee success across both genders AND 2. The AI system predicted that women would do worse than men That's politically embarrassing and something that you can't necessarily 'fix' by improving the system. (see: all the 'will this per…
One major problem. The parole software was NOT being fed data for "will this person commit another crime". It was being fed data for, "will this person be a suspect for another crime". The significant difference is that selective enforcement biases the data that it was trained on. Said selective enforcement has multiple causes, including the fact that heavier patrolling in black neighborhoods makes catching crimes mo…
Re: Notes on AI Bias
#118Earlier quoted context omitted.
What I was getting at is that our important choices are about outcomes and those have nothing really to do with assumptions about reality. For example all should be equal before the law. A statement that is supposed to be true but very obviously isn’t. Your post is great for the assumptions it encodes. Like what does it mean to be good at weight lifting? And that for some reason being good at weight lifting is a good…
> Like what does it mean to be good at weight lifting? And that for some reason being good at weight lifting is a good proxy for being a good bouncer or construction worker? being a good weight lifter, means you can lift heavier weights then a less-good weight lifter. Whether this is a proxy for anything isn't relevant, because it's a purely contrived example. There are clearly jobs where physical strength (among oth…
Re: Notes on AI Bias
#119Earlier quoted context omitted.
> Like what does it mean to be good at weight lifting? And that for some reason being good at weight lifting is a good proxy for being a good bouncer or construction worker? being a good weight lifter, means you can lift heavier weights then a less-good weight lifter. Whether this is a proxy for anything isn't relevant, because it's a purely contrived example. There are clearly jobs where physical strength (among oth…
Those were rhetorical questions I fully understand what the point of the example was.
"bias" and "reality" can equally cover for model simplicity.
Re: Notes on AI Bias
#120Earlier quoted context omitted.
Science has trained experts thinking about the data. If you set a team of scientists to find a way of predicting failure of turbines, they might notice a correlation between Siemens sensors and failure. They would then look for and attempt to prove theories to explain this descrepency. In doing so, they would likly discover that, not only can they not find a causative theory, but the correlation goes away when they c…
That's an interesting way to frame it. AI may stop at proximate causes rather than finding root causes