Just wait until people figure out that group composition can also affect whether people appear to be assholes, apathetic, etc. The wrong grouping will give even the most otherwise-perfect wallpower the permission and impetus to speak up and completely torch their opportunity.
On top of that, one evaluator in the right/wrong evaluation group can change the very definition of asshole on the spot and get group buy-in, given specific group composition.
To me a big part of the issue is that it's one test for one property, in one group configuration, by one evaluation group configuration...one might say it's _singularly_ disappointing to hear about that aspect.
One would hope a technology company could see the value in more scientific testing principles at least at a basic level? Hope my straw goggles are on, showing me straw-structures in straw-corporations.