Live data from Hacker News

New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

intextbooks.science.uu.nl

41–50 of 126 posts

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#41

Earlier quoted context omitted.

I would add that somewhere in there should be a spaced repetition algorithm. Spaced repetition is very effective, but it's really really clunky to use. My unpopular opinion is that we all have Stockholm syndrome when it comes to creating "cards", and people talk about how valuable creating cards is; but I think it stucks, it takes a lot of time. If AI is already teaching me math (let's say), it would be nice to tell…

But the very act of making and organizing your card deck is part of the SRS! It “sucks” because you get no dopamine hit from a fresh desk, as the reward system is not yet in place.

Again, I really think this is a viewpoint we've talked ourselves into to help us feel better about how cumbersome creating the cards are.

I'm willing to grant that there is some value in choosing what to put in the cards, but most of the awkwardness around making cards is UI related. Nobody creates cards on their phone, or while they're walking (AI could do both of these) - people create cards sitting at their computer (like cavemen!) usually clicking through a clunky UI and managing thousands of cards with thousands of clicks. That sucks, and people probably wont realize it sucks until something better comes along.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#42
The title is misleading. This isn't an AI tutor so much as a practice quiz platform with an AI autograder.

> constructed-response questions (CRQ) are graded by Claude Sonnet 4.6 against instructor-defined, question-specific rubric criteria

> Crucially, LLMs make it feasible to grade formative CRQ against rubric criteria at scale, a capability that appears pedagogically significant rather than merely convenient.

They specifically call out that the "RAG chat assistant" part of Phosphor (the platform) wasn't used much.

I commend the effort here, but I don't think these results are particularly noteworthy. The conclusion is essentially that people who do practice quizzes will do better on exams.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#43
post #10
post #2

Do you have a larger study planned for the Fall? It definitely seems promising. I'm curious how well you feel this worked because the subject was Statistics (objective grading) versus something more subjective like Civics or Literature. PS - I'd say this qualifies for Show HN, too! Do you

They were using Sonnet 4.6 for some fre form responses so that could be applied to something subjective.

But it's not clear that using Sonnet or any other LLM as a "grader" would result in the same improvement. For objective grading, you could be sure that the additional adaptive support is helping. For subjective things like writing style, literature, poetry, you end up with whatever Sonnet thinks is good (and randomly so).

It still could be better for students, but it's not obvious that it would be (or maybe not as strongly?).

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#44

I am somewhat skeptical of this. First, the headline result of 0.7*sigma improvement is the output of a statistical based on lessons/reviews they engaged with and their mid-term score, with that shift being for "full engagement". Based on their tables something like ~16 students (11% of the group) actually reached that level of engagement Second, trying to incorporate past grades into their modelling is not a substit…

This is a helpful explanation - am not a researcher so I have little idea how to run an unbiased, meaningful experiment (except that it takes a lot of effort and thought to run one). Useful analysis

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#45

Earlier quoted context omitted.

But the very act of making and organizing your card deck is part of the SRS! It “sucks” because you get no dopamine hit from a fresh desk, as the reward system is not yet in place.

Again, I really think this is a viewpoint we've talked ourselves into to help us feel better about how cumbersome creating the cards are. I'm willing to grant that there is some value in choosing what to put in the cards, but most of the awkwardness around making cards is UI related. Nobody creates cards on their phone, or while they're walking (AI could do both of these) - people create cards sitting at their comput…

> Nobody creates cards on their phone, or while they're walking

Wait, when are you doing it then? No wonder you think it sucks! Adopt some modern tools, yo. Use Anki, or vibe code your own app.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#46
This is super, but students will have access to AI during the test in real life, so it's ironically less realistic to remove it (thinking of the "... GPT-4 actually harmed subsequent performance by 17% when the tool was removed ..." part).

I'm more curious how students perform on the test with vs. without AI.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#47
post #42

The title is misleading. This isn't an AI tutor so much as a practice quiz platform with an AI autograder. > constructed-response questions (CRQ) are graded by Claude Sonnet 4.6 against instructor-defined, question-specific rubric criteria > Crucially, LLMs make it feasible to grade formative CRQ against rubric criteria at scale, a capability that appears pedagogically significant rather than merely convenient. They…

> a practice quiz platform with an AI autograder.

What do you think tutoring is?

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#48
post #14

Shocking that a well executed AI tutor improves outcomes. Hasn't computer assisted interactive learning already been proven for years? Why does there seem to be so much skepticism about enhancing it with AI? Is this just something like, astoundingly slow adoption or poor execution? Being held back by paper textbook makers? Teachers unions dragging their feet? How can interactive AI driven individually paced learning…

Lots of people in education will happily tell you how the past 15 years of tech integration has been a net negative. There ARE technologies that have improved things, but so much high-cost useless tech has been shoved into every level of education that many educators are incredibly leery of new tech. The issue is that while the underlying technology is useful, the way it gets integrated is frequently not. An administ…

I had a chance to use Google Classroom for a non-profit I was volunteering with, and wow it sucks. If that is what teachers and students have to contend with, yeah, I'd push back against any and all tech forced on me as well. It's all well intentioned, but the road to hell is paved with them.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#50
The article explicitly calls out selection bias (this is entirely based on 90% that opted into using the tutor, there was no control group), I wish the headline did as well. "Engaged students score 0.71 - 1.30 SD better in tests" sounds like a much simpler explanation.
Post reply on HN