Earlier quoted context omitted.
> Performance on the AT-SAT is not job performance. No, but it was the best predictor of job performance and academy pass rate there was. https://apps.dtic.mil/sti/pdfs/ADA566825.pdf https://www.faa.gov/sites/faa.gov/files/data_research/resear... (page 41) There are a fixed number of seats at the ATC academy in OKC, so it's critical to get the highest quality applicants possible to ensure that the pass rate is as hig…
That is NOT what the first study you've cited says at all: > "The empirically-keyed, response-option scored biodata scale demonstrated incremental validity over the computerized aptitude test battery in predicting scores representing the core technical skills of en route controllers." I.e the aptitude test battery is WORSE than the biodata scale. The second citation you offered merely notes that the AT-SAT battery is…
You're mistaken, it's the opposite. The first one found that AT-SAT performance was the best measure, with the biodata providing a small enhancement:
> AT-SAT scores accounted for 27% of variance in the criterion measure (β=0.520, adjusted R2=.271,p> In other words, after taking AT-SAT into account, CBAS accounted for just a bit more of the variance in the criterion measure
Hence, "incremental validity."
> The second citation you offered merely notes that the AT-SAT battery is a better predictor than the older OPM battery, not that is the best.
You're right, and I can't remember which study it was that explicitly said that it was the best measure. I'll post it here if I find it. However, given that each failed applicant costs the FAA hundreds of thousands of dollars, we can safely assume that there was no better measure readily available at the time, or it would have been used instead of the AT-SAT. Currently they use the ATSA instead of the AT-SAT, which is supposed to be a better predictor, and they're planning on replacing the AT-SAT in a year or two; it's an ongoing problem with ongoing research.
> I'd also say at a higher level that both of those papers absolutely reek of non-reproduceability and low N problems that plague social and psychological research. I'm not saying they're wrong. They are just not obviously definitive.
Given the limited number of controllers, this is going to be an issue in any study you find on the topic. You can only pull so many people off the boards to take these tests, so you're never going to have an enormous sample size.