Earlier quoted context omitted.
Do you also think LLM leaderboards accurately reflect the capabilities of the models being tested? If you do, then I can easily point you to numerous academic papers pointing out the various flaws in many leaderboards (from poorly designed benchmarks like bABI and the original SQuAD, to data contamination, and more). In that same way, any test, including the SAT and GRE have flaws. They can be gamed in ways similar t…
Do you think that LLM leaderboards don’t? Do you think a Llama 3 is going to beat an Opus 4.7 on any leaderboard? The real issue is that standardized tests disenfranchise lower SES students less than any other metric. Everyone who takes the SAT has to sit in the same room for the same amount of time answering the same questions. You can’t just pay someone else to take it for you (like essays) or select which difficul…
The issue is that the test is positively correlated with success in an undergraduate program, so they threw out the baby with the bathwster. The real issue is that the SAT is not able to distinguish the capabilities among students to the degree it purports to.