Live data from Hacker News

Flawed Algorithms Are Grading Millions of Students’ Essays

vice.com

281–290 of 302 posts

Re: Flawed Algorithms Are Grading Millions of Students’ Essays

#281
post #249
post #195

Earlier quoted context omitted.

> Turns out the instructions said the essay would be scored on verbal efficiency; getting the point across clearly with the fewest words. I started playing around and realized that the more words I added, the higher the score, whether they were relevant or grammatical or not. Frankly, this does not change anything from my experience in school decades ago. The teachers always said that the length does not matter and w…

Historically on the written portion of the SAT length is substantially correlated to the final score.

I'm not sure that's an issue by itself. If the prompt is broad enough, a minimum length can be reasonable for essays.

It certainly could be a problem if the prompt was too narrow, or time constraints, or some other factor.

Do you think the correlation by itself indicates something negative/inefficient I'm missing?

Re: Flawed Algorithms Are Grading Millions of Students’ Essays

#282

Earlier quoted context omitted.

Thus aren't you as well?

I meant that your self-consciousness and constant worry is probably harming your social interactions more than anything. Or not. It was for me in the past.

Aw, I was hoping we were going to dig deep into a "no u" situation.

Re: Flawed Algorithms Are Grading Millions of Students’ Essays

#283
post #168

Earlier quoted context omitted.

Depends upon what false negative rate you're willing to tolerate. ;) And I don't know how good of a signal there is. This is pure handwaving. But this type of thing seems like the exact kind of spooky correlation that ML is good at spotting compared to humans.

Machine learning techniques are going to be absolutely awful in detecting something like this, the reason being it's exceedingly rare (at least I'm guessing it is; if we're talking about child sexual abuse by one's own parents, it sure sounds extremely unlikely- but even child abuse in general is probably rare [1]). Machine learning systems are awful at identifying rare events. Like the OP seems to suggest, the false…

Humans are awful at rare events and vigilance tasks, too. That's part of why we're seeing machine vision and machine learning starting to outperform humans in e.g. grading radiology screening scans.

The total incidence of child abuse of all types from infancy to adulthood is on the order of 1 in 3. This is not terrifically rare-- it's of higher prevalence than pregnancy and of positive screening events.

A much bigger concern is non-causative correlations. It'd be pretty easy to train ML to be racist or look for e.g. indicators of class, which are correlates of abuse.

As to false positive rates-- you can pick your false positive rate to be whatever you want it to be, by twiddling the threshold for a positive result. I'm not sure false positives are of that great of a concern, if the output from a system is a notification to school administrators that they may want to keep an eye out for this student.

Re: Flawed Algorithms Are Grading Millions of Students’ Essays

#284

Earlier quoted context omitted.

Who cares? The people whose lives are ruined by being mis-identified by the system. The positives must be evaluated by a human anyway. Those same people whose lack of competence people are bemoaning throughout these comments.

> The people whose lives are ruined by being mis-identified by the system. When a child writes "daddy touches me between the legs" in an essay, it doesn't matter if a human spots it or an AI that forwards it to a human, this needs to be investigated either way. > Those same people whose lack of competence people are bemoaning throughout these comments. It's not a lack of competence that's bemoaned, it's a massive amo…

> When a child writes "daddy touches me between the legs" in an essay, it doesn't matter if a human spots it or an AI that forwards it to a human, this needs to be investigated either way.

When a child writes a set of things that individually are not very concerning, they may have cues that could say "hey, this kid, you should maybe keep an eye out for evidence of abuse."

Particularly attuned, experienced individuals might spot these cumulative cues, but we all know that this is not all people dealing with children.

It's an interesting problem.

Re: Flawed Algorithms Are Grading Millions of Students’ Essays

#285
post #164

Earlier quoted context omitted.

so ... nobody wonders about the obvious rammification: then any ML scoring systems ... must detect child abuse signals!

Hey uh, that actually seems valuable. I'd believe that ML could spot abuse that humans miss pretty well from signals like non-overt references in homework and school records, if one could come up with an adequate training set. Much more likely than teaching ML to score reasoned and creative activity in any reasonable way.

This is so dangerous.

Society's bigotry is going to flood that bad boy so quick you might as well name it Gobbels.

I love ML. I want children to be safe. This is not the place for ML or AI or Quantum or any tech.

What needs to exist is better resources for those children, that mother grading the tests, the teachers of those children, and social services that are meant to support them. If you want to make a difference about this, look there.

Don't go building a automaton King Solomon who decides why this kid should be taken from these parents because speaking Spanish was worth -0.1 on some goddamn weight trained on data generated from a racist society.

This isn't a "spooky" correlation a cool algorithm can detect, it's a serious, layered social problem.

Re: Flawed Algorithms Are Grading Millions of Students’ Essays

#286
post #168

Earlier quoted context omitted.

What do you think the false positive rate is likely to be?

Depends upon what false negative rate you're willing to tolerate. ;) And I don't know how good of a signal there is. This is pure handwaving. But this type of thing seems like the exact kind of spooky correlation that ML is good at spotting compared to humans.

> But this type of thing seems like the exact kind of spooky correlation that ML is good at spotting compared to humans.

How? Particularly, where do you get training data at the required scale?

Re: Flawed Algorithms Are Grading Millions of Students’ Essays

#287

Earlier quoted context omitted.

Why not? The school contracts with the testing agency that does the grading. Seems like a contractor relationship?

School contractor is someone who is hired by the school district on via a contractual relationship. Think temporary teachers, or custodian staff. It’s not a transitive relationship to every employee of every company who has some sort of contract, however small, with a school.

every relationship however transitive or small, is a relationship too! (Dr. Seuss)

Re: Flawed Algorithms Are Grading Millions of Students’ Essays

#288
post #168

Earlier quoted context omitted.

Depends upon what false negative rate you're willing to tolerate. ;) And I don't know how good of a signal there is. This is pure handwaving. But this type of thing seems like the exact kind of spooky correlation that ML is good at spotting compared to humans.

Machine learning techniques are going to be absolutely awful in detecting something like this, the reason being it's exceedingly rare (at least I'm guessing it is; if we're talking about child sexual abuse by one's own parents, it sure sounds extremely unlikely- but even child abuse in general is probably rare [1]). Machine learning systems are awful at identifying rare events. Like the OP seems to suggest, the false…

> if we're talking about child sexual abuse by one's own parents, it sure sounds extremely unlikely-

Child sexual abuse isn't extremely rare and familial abuse is a very large minority of child sex abuse.

Re: Flawed Algorithms Are Grading Millions of Students’ Essays

#289

Earlier quoted context omitted.

> I feel like you'd probably have to have an AGI to meaningfully evaluate an essay. So the reason this isn't the case, is because there are very simple metrics that tend to highly correlate with essay quality. It doesn't mean the grading-bot is actually evaluating essay quality. It's just looking for properties that are statistically associated with good essays. Remember, at the end of the day as long as the bot's ra…

I can't disagree strongly enough. >Remember, at the end of the day as long as the bot's ranking is close enough to the human grader's ranking, nobody really cares about the internal logic. This isn't true at all. Imagine you got a B or C on an essay that a human would have given an A to because you wrote it concisely and in plain language, or because you used language that's statistically correlated with being black.…

Minor correction; The automated resume reviewer was biased against women according to your reference.

Re: Flawed Algorithms Are Grading Millions of Students’ Essays

#290
post #168

Earlier quoted context omitted.

Depends upon what false negative rate you're willing to tolerate. ;) And I don't know how good of a signal there is. This is pure handwaving. But this type of thing seems like the exact kind of spooky correlation that ML is good at spotting compared to humans.

> But this type of thing seems like the exact kind of spooky correlation that ML is good at spotting compared to humans. How? Particularly, where do you get training data at the required scale?

You take samples of hundreds or thousands of past students' schoolwork, e.g. submissions of essays for standardized tests.

You survey those kids in adulthood about whether and how they were subject to abuse and other types of relevant adversity.

You attempt to control the data so that you don't just latch onto other correlates of abuse (e.g. social class).

Post reply on HN