Live data from Hacker News

New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

intextbooks.science.uu.nl

51–60 of 126 posts

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#51
While there's some skepticism in the thread, I'm not particularly surprised if this is true. Children who can get human tutoring do a lot better. An LLM that can answer questions and patiently explain likely offers some benefit.

What creeps me out about bringing LLM into early education is that it's a period where kids learn to socialize and cope with problems, and I do worry about forming substitute relationships with chatbots that are engineered for sycophancy / enablement. But I guess that's a problem either way, because almost every student will try an LLM at some point.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#52
post #42

The title is misleading. This isn't an AI tutor so much as a practice quiz platform with an AI autograder. > constructed-response questions (CRQ) are graded by Claude Sonnet 4.6 against instructor-defined, question-specific rubric criteria > Crucially, LLMs make it feasible to grade formative CRQ against rubric criteria at scale, a capability that appears pedagogically significant rather than merely convenient. They…

> a practice quiz platform with an AI autograder. What do you think tutoring is?

Definitely not just grading. Tutoring is explaining and back and forth discussion to impart knowledge, in context and in response to specific difficulties/confusion the student is having.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#53
Interesting article, wonder where we're going with this though, I find it's very difficult to keep LLMs on track and critical enough to be useful.

Just want to say that:

>In our deployment, student-reported reading completion baselines for MATH 010 were approximately 15%, with instructors estimating 10%. Individual student reports of reading compliance ranged from "literally no one does that" to "is this being recorded?"

is hilarious

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#54
post #17

Conflicted about this study. On one hand, LLMs have been incredible for my personal learnings of new concepts. On the other, I'm sceptical of that it'll have "strong benefits" at scale; I'd be more in favor if the wording was "some"/"moderate". I reckon self-selection plays a huge part, as mentioned in the "Limitations" section of the paper. I'd also caution against attaching the tool to grading. That means students…

> LLMs have been incredible for my personal learnings of new concepts. Mind if I ask what did you learn and how you're using it? The reason I'm asking is that I repeatedly felt excitement only to realize down the line that the explanations didn't actually translate into practical skills. I'm not sure it's even an AI problem, it's a "doing versus reading" problem. Same as with reading a pop-science article and thinkin…

Various concepts when I joined new teams in domains I've never worked in before. And system design. So very practical, and where stakes were high.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#55
post #50

The article explicitly calls out selection bias (this is entirely based on 90% that opted into using the tutor, there was no control group), I wish the headline did as well. "Engaged students score 0.71 - 1.30 SD better in tests" sounds like a much simpler explanation.

I used to TA a graduate level CS math class at Georgia Tech. We regularly saw that the students who self-organized study groups did dramatically better in the course than average. One semester they told us to put everyone in study groups to see if it helped. The effect disappeared. Turns out that it was the self-selection of the most engaged students into a small group that mattered, not the study group itself.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#57
post #11

I'm on record saying that a system like this with some extra hardware (i.e. a way for the LLM to have live understanding of the student's paper notebook or handout which are being written in with a plain old pencil) combines the best of both worlds - individual tutoring with approximately zero screen time which scales linearly with the number of students. The role of the teacher or professor then becomes a manager of…

Today I saw a demo of Remarkable turned into Voldemort's diary from Harry Potter - you write to it, and it writes back, in handwriting.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#58

I am somewhat skeptical of this. First, the headline result of 0.7*sigma improvement is the output of a statistical based on lessons/reviews they engaged with and their mid-term score, with that shift being for "full engagement". Based on their tables something like ~16 students (11% of the group) actually reached that level of engagement Second, trying to incorporate past grades into their modelling is not a substit…

Worse, because students complained about the difficulty of the AI-graded quizzes, they switch to multiple-choice questions only, which increases engagement, but after analyzing the exam results they determine that multiple-choice questions don't seem to help and add AI-graded questions back, after which engagement drops again.

That means their experiment design is partially caused by their results instead of the other way around, which is a bad situation to be in. Their statistical analysis is completely inadequate for dealing with this.

And the change in engagement suggests that there's strong selection involved. Their attempt to use midterm scores to control for selection effects is unconvincing. Why not control for whether students used the platform more when there were only multiple-choice questions? Those are the ones who self-selected out of using the AI grader.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#59
A lot of pessimism in the comments, but I am just happy that we are seeing some work towards bridging the 2 Sigma gap for regular education vs. elite private tutoring. I can't imagine that people assume it's the physical presence of the tutor that is making the difference, it has to come down to the personalisation and expertise which is exactly what AI can provide in a form. And yea it might not be "there" yet. But if we don't start trying and studying then it'll never get there.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#60

A lot of pessimism in the comments, but I am just happy that we are seeing some work towards bridging the 2 Sigma gap for regular education vs. elite private tutoring. I can't imagine that people assume it's the physical presence of the tutor that is making the difference, it has to come down to the personalisation and expertise which is exactly what AI can provide in a form. And yea it might not be "there" yet. But…

Tell me if I am oversimplifying, but I never understood the noise about the two sigma problem. Like, of course if you have a private tutor to immediately answer any question that pops into your head at the immediate moment you get confused, you are going to learn vastly more efficiently than in a large classroom where once you get confused you are likely to stay confused. To say nothing of how the pace will likely either drag way behind what you'd like, or accelerate too fast ahead of it.

The environment is just obviously two sigma better. This just... seems obvious to me? In the same way that I will get stronger much faster if I have a physical trainer to tell me exactly what I am doing wrong when I do it? And it seems obviously unsolvable other than by getting everyone a private tutor (or AI..?).

Asking from a place of curiosity.

Post reply on HN