1. Give a sample of mechanical turk workers the same data to classify without telling them
2. Once submitted, look carefully at the ones they get wrong
3. If the turk worker gets similar ones wrong to other turk workers, they're probably not cheating
4. If the turk worker is frequently wrong on questions other turk workers all got right (e.g. if 99.9% of other workers got it right) then they've likely used AI to generate their answers and then fiddled around with a few afterwards to make it look human
That presumes they're fiddling with a few responses for the human touch. If they're submitting raw AI output that should be easily discoverable by running similar models against the same data (like detection of chess cheating compares games against moves form several engines, not just one).
I think that would work, unless everybody uses AI. If that latter were suspected, get a group to label data in person (where you can be 100% sure they can't cheat), then use those answers as the basis for comparison.