Earlier quoted context omitted.
I agree there is likely impact from the recognition itself. Shouldn't be hard to verify, no? Have him pick 2N candidates, randomly select half that he can not announce publicly, talk to or otherwise influence (impossible to prevent entirely but "good enough" will suffice to establish a difference). Then look at how the two groups are doing some time down the line.
Not hard to verify in some weird vacuum where completely pointless albeit simple sounding things are done to validate random conjectures, yes.
You can't improve what you can't measure.
But is the goal of these sorts of things often actually to select the best?