Earlier quoted context omitted.
The comparison in the paper is who approximates trained annotators better, MTurk or ChatGPT. Trained annotators are the gold standard.
The "gold standard" in the article is two graduate students. The Mechanical Turks are probably more- I can't find the number in the paper. Given the "inter-coder agreement" is low for the Mechanical Turks (0.17% Pearson cor. coeff) it is no surprise that the "gold standard" and the Mechanical Turks' decisions diverge. So the comparison is pointless.
It seems to me that, A, at least for some tasks, grad-students-as-gold-standard is not wacky, and B, grad students won't necessarily have the same output as MTurk workers. Given these two points, it's perfectly reasonable to ask how consistent-with-grad-students MTurk workers are, seeing as the reason people use MTurk in similar scenarios is as a cheaper alternative to grad students.
Re: inter-coder agreement, I think that the low inter-coder agreement for MTurk is in itself potentially a surprising aspect. Perhaps the explanations that worked well for grad students (and ChatGPT) didn't explain the task properly to MTurks. That's pretty much a point in favor of ChatGPT, though maybe it can be viewed as a point against the actual tasks they use (ie if the explanation couldn't get people of a likely more dissimilar background to output the same results as grad students then maybe the task is less objective than deemed by the task's authors).
Side-notes: - I looked for the 0.17% number in the paper, and what the paper actually says is that the Pearson correlation coefficient between inter-coder agreement and classification accuracy is 0.17 (=positive but weak). This isn't a number comparing the correlation between different coders. (Frankly, I think comparing ChatGPT's self-consistency to consistency between different people is unenlightening) - Section 4 sub-section "Crowd-Workers Annotation" says that each tweet was classified by two different workers, and no single worker classified more than 20% of the dataset.