Live data from Hacker News

New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

intextbooks.science.uu.nl

11–20 of 126 posts

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#11
I'm on record saying that a system like this with some extra hardware (i.e. a way for the LLM to have live understanding of the student's paper notebook or handout which are being written in with a plain old pencil) combines the best of both worlds - individual tutoring with approximately zero screen time which scales linearly with the number of students. The role of the teacher or professor then becomes a manager of the student - agentic tutor pairs, a referee when the student and model disagree, etc. and most importantly still being the human teacher you can just talk to in the human education process.

I'm convinced this is the future of education - models are there, we need the classroom tech to catch up. The alternative is obvious and quantified in the paper - students just use models to do their work for them and learn nothing.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#12

Honestly whether or not this was effective seems less important to me than the adoption numbers. Text book reading in this course was 10-15% at baseline ... but this AI thing got 90% voluntary usage ungraded. Even if its worse per-hour than a textbook, you're now teaching 6x as many students _something_ instead of teaching a small minority everything. So really it just becomes an optimization problem at that point be…

I'd argue the results are even better: just reading a textbook doesn't really teach you much. You have to do exercises, but they're expensive to create and grade. LLMs with a proper harness (see paper) tackle both.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#13
post #4

Too bad the educational use case doesn't make any money. Good LLMs are a game changer for people motivated to learn.

Wikipedia doesn’t make much money but is still helpful. LLMs don’t need to make a whole bunch of money to be helpful.

People aren't paying trillions to train them to be helpful. They want to make quadrillions.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#14
Shocking that a well executed AI tutor improves outcomes.

Hasn't computer assisted interactive learning already been proven for years? Why does there seem to be so much skepticism about enhancing it with AI?

Is this just something like, astoundingly slow adoption or poor execution? Being held back by paper textbook makers? Teachers unions dragging their feet?

How can interactive AI driven individually paced learning _not_ be obviously dramatically more effective?

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#15
post #7
post #4

Too bad the educational use case doesn't make any money. Good LLMs are a game changer for people motivated to learn.

I don't want to learn from hallucinations where it will change its answers based on me questioning their teachings. I use it for conversations in a language I'm learning, but I quickly learned that asking it grammar questions for example is not a wise decision.

Curious whether you were just bare asking it questions, or whether you provided it with lessons one by one with instruction that the lesson is the baseline truth etc

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#16
post #14

Shocking that a well executed AI tutor improves outcomes. Hasn't computer assisted interactive learning already been proven for years? Why does there seem to be so much skepticism about enhancing it with AI? Is this just something like, astoundingly slow adoption or poor execution? Being held back by paper textbook makers? Teachers unions dragging their feet? How can interactive AI driven individually paced learning…

its like anything else. benifits students that are already motivated to learn.

very few are actually motivated to learn and are just there to get a job or its just next thing that they have to do in life.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#17
Conflicted about this study. On one hand, LLMs have been incredible for my personal learnings of new concepts.

On the other, I'm sceptical of that it'll have "strong benefits" at scale; I'd be more in favor if the wording was "some"/"moderate". I reckon self-selection plays a huge part, as mentioned in the "Limitations" section of the paper.

I'd also caution against attaching the tool to grading. That means students have to put more effort into the course, which increases the chances that they will use LLMs to save time rather than make the investment.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#18
post #7
post #4

Too bad the educational use case doesn't make any money. Good LLMs are a game changer for people motivated to learn.

I don't want to learn from hallucinations where it will change its answers based on me questioning their teachings. I use it for conversations in a language I'm learning, but I quickly learned that asking it grammar questions for example is not a wise decision.

Are we talking about human teachers or LLMs here?

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#19

Honestly whether or not this was effective seems less important to me than the adoption numbers. Text book reading in this course was 10-15% at baseline ... but this AI thing got 90% voluntary usage ungraded. Even if its worse per-hour than a textbook, you're now teaching 6x as many students _something_ instead of teaching a small minority everything. So really it just becomes an optimization problem at that point be…

[deleted]

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#20
post #6

This is exciting because the effect size is so large. But as the author's acknowledged, selection bias is nearly impossible to control for in this non-randomized study: > and lacks randomized controls. Self-selection is the central threat: students who complete more quizzes may be more motivated or higher-performing generally But this is still a strong result. I'm excited to see more in this space.

They tried to control for this. It's described in the first paragraph of section 4.
Post reply on HN