Earlier quoted context omitted.
The problem is that metric, like any metric, is both easy to game _and_ can provide misleading information. Measuring number of commits? Create fewer, larger commits. Measuring commit size? Pull in more third-party libraries, even where it does't make sense. Author count? Add more/less documentation and recruit or inhibit new devs depending on what your goal is. Not to mention the number of commits/authors before and…
Ok, then find another way to measure developer productivity, or reliability in production, or customer features delivered. If you can’t find a measurable benefit to a refactoring (or anything else, really) then maybe it was not worth doing in the first place.
There is no programming project in existence with enough developers working on it, that developer-productivity data derived from a change to it would not be considered "underpowered" for the sake of proving anything.