What creeps me out about bringing LLM into early education is that it's a period where kids learn to socialize and cope with problems, and I do worry about forming substitute relationships with chatbots that are engineered for sycophancy / enablement. But I guess that's a problem either way, because almost every student will try an LLM at some point.
New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
51–60 of 126 posts
Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#52The title is misleading. This isn't an AI tutor so much as a practice quiz platform with an AI autograder. > constructed-response questions (CRQ) are graded by Claude Sonnet 4.6 against instructor-defined, question-specific rubric criteria > Crucially, LLMs make it feasible to grade formative CRQ against rubric criteria at scale, a capability that appears pedagogically significant rather than merely convenient. They…
> a practice quiz platform with an AI autograder. What do you think tutoring is?
Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#53Just want to say that:
>In our deployment, student-reported reading completion baselines for MATH 010 were approximately 15%, with instructors estimating 10%. Individual student reports of reading compliance ranged from "literally no one does that" to "is this being recorded?"
is hilarious
Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#54Conflicted about this study. On one hand, LLMs have been incredible for my personal learnings of new concepts. On the other, I'm sceptical of that it'll have "strong benefits" at scale; I'd be more in favor if the wording was "some"/"moderate". I reckon self-selection plays a huge part, as mentioned in the "Limitations" section of the paper. I'd also caution against attaching the tool to grading. That means students…
> LLMs have been incredible for my personal learnings of new concepts. Mind if I ask what did you learn and how you're using it? The reason I'm asking is that I repeatedly felt excitement only to realize down the line that the explanations didn't actually translate into practical skills. I'm not sure it's even an AI problem, it's a "doing versus reading" problem. Same as with reading a pop-science article and thinkin…
Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#55The article explicitly calls out selection bias (this is entirely based on 90% that opted into using the tutor, there was no control group), I wish the headline did as well. "Engaged students score 0.71 - 1.30 SD better in tests" sounds like a much simpler explanation.
Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#56Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#57I'm on record saying that a system like this with some extra hardware (i.e. a way for the LLM to have live understanding of the student's paper notebook or handout which are being written in with a plain old pencil) combines the best of both worlds - individual tutoring with approximately zero screen time which scales linearly with the number of students. The role of the teacher or professor then becomes a manager of…
Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#58I am somewhat skeptical of this. First, the headline result of 0.7*sigma improvement is the output of a statistical based on lessons/reviews they engaged with and their mid-term score, with that shift being for "full engagement". Based on their tables something like ~16 students (11% of the group) actually reached that level of engagement Second, trying to incorporate past grades into their modelling is not a substit…
That means their experiment design is partially caused by their results instead of the other way around, which is a bad situation to be in. Their statistical analysis is completely inadequate for dealing with this.
And the change in engagement suggests that there's strong selection involved. Their attempt to use midterm scores to control for selection effects is unconvincing. Why not control for whether students used the platform more when there were only multiple-choice questions? Those are the ones who self-selected out of using the AI grader.
Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#59Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#60A lot of pessimism in the comments, but I am just happy that we are seeing some work towards bridging the 2 Sigma gap for regular education vs. elite private tutoring. I can't imagine that people assume it's the physical presence of the tutor that is making the difference, it has to come down to the personalisation and expertise which is exactly what AI can provide in a form. And yea it might not be "there" yet. But…
The environment is just obviously two sigma better. This just... seems obvious to me? In the same way that I will get stronger much faster if I have a physical trainer to tell me exactly what I am doing wrong when I do it? And it seems obviously unsolvable other than by getting everyone a private tutor (or AI..?).
Asking from a place of curiosity.