As someone who works in this field: the challenge is more nuanced. For example: Even if you don’t input ethnicity as a value, you might input ZIP codes. ZIP codes however can be very predictive of someone’s ethnic background, and you might still end up unfairly discriminating against certain groups of people.
The Amazon recruiting AI failure is another great example of this, I think it didn’t take gender into account, but discriminated against activities like ‘head of women’s association’ etc.
Finally, what’s ethical is often up for discussion. Depending on context, people will optimize for equality (equal % of credits approved for both genders) vs helping those behind (Higher % approved for those historically disadvantaged) vs helping those that ‘deserve’ it of groups were unknown (finance such that default rates of both groups are expected to be equal, likely benefiting those with already good credit scores). I don’t think any of these answers is wrong, but they are mutually exclusive.
One challenge I don’t see discussed that often is that the current system of many humans evaluating loan applications leads to ‘noise‘ in outcomes so that we (on average) don’t see that much bias. (Eg a few people of certain ethnicity not getting loans because a few people are biased against them)
A single AI system scoring all applications however is consistent in its bias against a group, leading to more obvious effects (eg most people in a certain area not receiving loans because the algorithm is biased against it)
In my opinion, an important part is the right on human review of meaningful automated decisions. Not mentioned that often, but it’s also part of the GDPR.