Earlier quoted context omitted.
I've started referring to this phenomena as "Goodhart's Hell." We're living in a world with quickly increasing complexity. Variance in samples is much larger because our reach in what we sample is larger (we talk to people far from our homes, with different cultures, interpretations of words, language patterns, behaviors, ethics, etc). But we've responded to this only by simplifying our evaluations. We've created mor…
In part it's due to requirements of transparency. If some stranger is going to be evaluating you, a simple way to ensure that you're not being evaluated unfairly is to have the evaluation criteria be very simple, almost to the point of being mechanistic. If you're part of a small enough community it's likely that your evaluator, their overseer, and your peers all know you, so it's very difficult for unfairness to cre…
I actually think scale is part of the problem here. In small communities there is more natural nuance and correction as there's high accountability. But with scale accountability often disappear as you disappear into the crowd. And you'll always be able to point at winners and losers to whichever noise you favor. The problem is we're not looking at things holistically and ensuring that our metrics are aligned with our desired outcomes. Instead we often pick a metric because it reasonably looks like it should align with the desired outcome (not proven, just intuited) and then just tell anyone that disagrees that it is obvious instead of acknowledging the downsides. Every metric has limitations. Ignoring them is the problem.