The main article tokenadult is talking about is the Schmidt & Hunter paper: "The conclusions in this article apply mainly to the middle 62% of jobs in the U.S. economy in terms of complexity. [...] This category includes skilled blue collar jobs and mid-level white collar jobs, such as upper level clerical and lower level administrative jobs."
The article says that the most common measure of employee ability in general, and presumably for the articles used in this meta-study, was the amount of money each employee earned. Presumably then most of these employees were doing routinized or semi-routinized labor that would lend itself to piecewise compensation. The above quote seems to be consistent with this interpretation. They are also only looking at employees of huge corporations, where each person is one of dozens or hundreds of others doing more or less the same thing.
So is this meta-study relevant to hiring programmers? Yes and no. Probably the measures that were found to have more validity are still going to have more validity than the ones that were found to have less validity. But at the same time using only these tests would be too simplistic; they were designed to predict which factory workers were likely to steal from the company or slack off. And while these are still important factors to consider, the challenges of a modern startup go way above and beyond this.
Essentially this research was designed mainly for situations where you're trying to scale up a large industrial process where you can make money by arbitraging the difference between the output of the average employee and the amount it costs to pay them. The closer you get toward environments where it's essential that each person contribute things that are unique and novel, the less it makes sense to rely on these sorts of hiring tools.
In short, I would say that the above research neither supports nor disconfirms the advice of the original blog post. Of course you could make the case (and maybe tokenadult believes this?) that employers should make hiring decisions mainly by using a series of multiple choice tests that have been empirically validated, but that's an entirely different argument.