It would make sense to use the AI as a first pass, and then not randomly grade the essays with a human, but specifically choose all the essays that are on the cusp of the pass fail line. Then use all those human generated scores to update the model, especially if someone moves from pass to fail or fail to pass. Then maybe throw in a few of the really high and really low outliers to make sure those are right, and throw away your entire model if the human scores are drastically different (and obviously don't tell the humans what the computer score was so they have no idea if they're reading a "cusp" essay or an outlier essay).
But putting the educational fate (and therefore future earnings) in the hands of an AI is unconscionable.