Two notes:
1. I don't see anything about the statistical significance, or the natural fluctuation of the results, i.e. the standard error. Presumably lots of students are measured, and you could argue that your standard error (proportional to 1/sqrt(N)) is tiny, but then, the questions are different every year, so you have some date-specific (not student-specific) error. In other words, just looking at previous years, how much do results vary? And is this year's flatline maybe just a result of measurement error? (In other words: Q: "Why are they flatlining?" A: No reason.)
EDIT to add: well, there is one sentence in the article:
> “not a significant change,” spokeswoman Amber Farinha wrote in an email
2. It's a
> computer adaptive test — meaning as the test progresses, questions become harder or easier depending on the student’s answers
These have problems, in my view. I had a Spanish girl friend that studied for the GRE with me. The verbal section is largely a long vocabulary test, so we studied quite a bit (from the entertaining Princeton Review Word Smart book). Now, often she would know the "complicated" word being explained, as it had Latin roots and was the same in Spanish. However, she wouldn't know the "easy" explanation. (For example, "to lament" (ah, easy, like "lamentablemente, lamentar") was explained as "to mourn" (what?), "arboreal" is "arbóreo" (well, "tree dweller" is easy, too, but you get the idea).)
Bottom line: we did several (actual old) practice exams, timed, and she would consistently score in a certain percentile range (with some variation, of course). Those were good old-fashioned paper tests, though.
However, when she did the actual (adaptive) GRE, she scored much worse than in the practice exams, several standard deviations away. (And before you blame it on nervousness due to real test conditions, note that there was a significant drop only in the verbal part, not analytic or quantitative.)
This appears to me a problem with the testing methodology: when she initially got something wrong, the CAT would give her "easier" questions, preponderantly words with Germanic roots, which she would also get wrong. The "hard" questions involving words with Latin roots, that she'd probably have gotten right, were never displayed to her.
She got into a good school anyway, so it didn't really matter. But in my view, the claim that the CAT is just as good as the longer paper based test is predicated on assuming that test subjects come from the same population. I really wonder whether speakers of Romance languages where systematically disadvantaged by the switch from paper to CAT, and whether ETS (the producer of GRE) did any systematic study on that.
(Note that it is conceivable that some Hispanophones got questions right at the beginning, and then got "harder" questions easier for them, thus getting better results than in the paper based test. If that is the case, it's conceivable that the average scores of both the Anglo- and Hispanophone group lines up with the non-adaptive and adaptive test, but Hispanophones suddenly show a much bigger variance in the adaptive test. Questions questions.)