While there's some skepticism in the thread, I'm not particularly surprised if this is true. Children who can get human tutoring do a lot better. An LLM that can answer questions and patiently explain likely offers some benefit. What creeps me out about bringing LLM into early education is that it's a period where kids learn to socialize and cope with problems, and I do worry about forming substitute relationships wi…
New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
61–70 of 126 posts
Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#62The article explicitly calls out selection bias (this is entirely based on 90% that opted into using the tutor, there was no control group), I wish the headline did as well. "Engaged students score 0.71 - 1.30 SD better in tests" sounds like a much simpler explanation.
This sentence is accurate, but inevitably leads to the confusion you see in these comments.
Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#63Earlier quoted context omitted.
A 'smart pen' that records the student's writing in some way, maybe? My first thought was a tablet that boots straight into a writing software but students should not be subjected to any amount of latency in their writing. Practically, I think if you want the AI system to have a live view of what the student's doing you're going to have to replace one of either the tablet or the writing instrument. A wearable camera…
there was a pen that used special paper to directly record your notes (15-20years ago)... should be possible nowadays to directly transfer this to a connected device and have it feed it to an llm. and after looking it up, it appears they are still available: https://www.livescribe.com/landingpage/ls3_onenote/
Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#64The article explicitly calls out selection bias (this is entirely based on 90% that opted into using the tutor, there was no control group), I wish the headline did as well. "Engaged students score 0.71 - 1.30 SD better in tests" sounds like a much simpler explanation.
I used to TA a graduate level CS math class at Georgia Tech. We regularly saw that the students who self-organized study groups did dramatically better in the course than average. One semester they told us to put everyone in study groups to see if it helped. The effect disappeared. Turns out that it was the self-selection of the most engaged students into a small group that mattered, not the study group itself.
If it's purely a correlation, then maybe those students would be more successful than average even without the study group. They're already the most motivated kids. Maybe they just do "motivated kid stuff" and would still outperform.
Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#65The title is misleading. This isn't an AI tutor so much as a practice quiz platform with an AI autograder. > constructed-response questions (CRQ) are graded by Claude Sonnet 4.6 against instructor-defined, question-specific rubric criteria > Crucially, LLMs make it feasible to grade formative CRQ against rubric criteria at scale, a capability that appears pedagogically significant rather than merely convenient. They…
> a practice quiz platform with an AI autograder. What do you think tutoring is?
Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#66Shocking that a well executed AI tutor improves outcomes. Hasn't computer assisted interactive learning already been proven for years? Why does there seem to be so much skepticism about enhancing it with AI? Is this just something like, astoundingly slow adoption or poor execution? Being held back by paper textbook makers? Teachers unions dragging their feet? How can interactive AI driven individually paced learning…
its like anything else. benifits students that are already motivated to learn. very few are actually motivated to learn and are just there to get a job or its just next thing that they have to do in life.
Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#67Earlier quoted context omitted.
I don't want to learn from hallucinations where it will change its answers based on me questioning their teachings. I use it for conversations in a language I'm learning, but I quickly learned that asking it grammar questions for example is not a wise decision.
Are we talking about human teachers or LLMs here?
Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#68I am somewhat skeptical of this. First, the headline result of 0.7*sigma improvement is the output of a statistical based on lessons/reviews they engaged with and their mid-term score, with that shift being for "full engagement". Based on their tables something like ~16 students (11% of the group) actually reached that level of engagement Second, trying to incorporate past grades into their modelling is not a substit…
Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#69Yes! Very exciting to see this. Bloom's Two Sigma Opportunity suggests that there's another SD improvement available: https://en.wikipedia.org/wiki/Bloom%27s_2_sigma_problem
The story around Bloom's two sigma is a bit complex https://nintil.com/bloom-sigma/
Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#70I'm on record saying that a system like this with some extra hardware (i.e. a way for the LLM to have live understanding of the student's paper notebook or handout which are being written in with a plain old pencil) combines the best of both worlds - individual tutoring with approximately zero screen time which scales linearly with the number of students. The role of the teacher or professor then becomes a manager of…
I work in consulting and one of my projects is piloting an AI use case for a department within one of my clients. On a discovery call someone casually brought up that they bought a reMarkable notebook themselves and were wondering if it could be integrated into the use case. It really got me thinking. Maybe reMarkable or something like it could help bridge a student's writing with an LLM without having to fall back t…
What does “bridge a student’s writing” mean?? If this is a real argument it needs to be clearer.
What’s the functional difference between a Remarkable and an iPad? The former is less responsive, costs less, and has better battery life, right? I really don’t see how that’s significant to any kind of development of anything.
Are you talking about running a local model??