Live data from Hacker News

New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

intextbooks.science.uu.nl

111–120 of 126 posts

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#111

Earlier quoted context omitted.

The tricky question at the forefront of education research is, at least in my mind, trying to thread the needle between effective techniques that students don’t like, ineffective techniques that students do like, and poorly defined techniques that administrators like. And on top of all that student self-reports aren’t actually very reliable indicators of learning progress at all!

>effective techniques that students don’t like Do students not like Mastership techniques?

Students don't like being told they're bad at something. Students don't like being told to do the thing they're bad at. They especially don't like having to constantly repeat this in a loop until they've caught up with the rest of the class, while other students get to enjoy their free time.

Also, mastery learning doesn't look all that effective on standardized tests: https://www.jstor.org/stable/1170613

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#112

Earlier quoted context omitted.

It feels to me that the venn diagram between "students that fully engaged with the material" and "students that learned well from the material" is going to basically be a circle for any teaching method.

The tricky question at the forefront of education research is, at least in my mind, trying to thread the needle between effective techniques that students don’t like, ineffective techniques that students do like, and poorly defined techniques that administrators like. And on top of all that student self-reports aren’t actually very reliable indicators of learning progress at all!

This first sentence is the best summary of educational problems I've ever read, thank you.

I think we've all had a few teachers who seem able to teach 2-10x more effectively than normal; they all seem to do it in different ways. I read recently the Gates Foundation felt they'd found nothing statistically significant in their billion+ dollars of teaching philanthropy.

The most plausible education improvement proposal I'm aware of is individual tutoring. This truly does seem to work well, and I think it's why there's so much interest in agentic tutoring. Perhaps we need a TeachBench to get some hill climbing done by the frontier labs.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#113

I am somewhat skeptical of this. First, the headline result of 0.7*sigma improvement is the output of a statistical based on lessons/reviews they engaged with and their mid-term score, with that shift being for "full engagement". Based on their tables something like ~16 students (11% of the group) actually reached that level of engagement Second, trying to incorporate past grades into their modelling is not a substit…

Thank you for the feedback! Maybe the following info will be helpful when considering our results: 1. Quiz completion is our deliberately conservative lower bound on reading compliance, and the 0.71 figure is not a claim that those 16 students each gained that much. The estimate is from a regression carried by the per-lesson slope, fit across the whole dosage distribution, and the underlying dosage-performance relati…

Thanks for this. No idea why you're getting downvoted on this background.

Anything promising that helps students learn and learn more effectively is obviously incredibly valuable. My concern here is that I don't read anything trying to disambiguate students who just, you know, work harder, and therefore used whatever materials they had to hand including Phosphor, from those who work at the same level and chose not to use Phosphor.

We all recall from our days in undergrad that there are students who do nothing and slide by, a few who do nothing and ace everything, and students who outwork everyone else and outperform -- I think a tool that only helps grinders who were already going to grind is likely not the contribution you're hoping to make.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#114
post #38
post #11

I'm on record saying that a system like this with some extra hardware (i.e. a way for the LLM to have live understanding of the student's paper notebook or handout which are being written in with a plain old pencil) combines the best of both worlds - individual tutoring with approximately zero screen time which scales linearly with the number of students. The role of the teacher or professor then becomes a manager of…

I work in consulting and one of my projects is piloting an AI use case for a department within one of my clients. On a discovery call someone casually brought up that they bought a reMarkable notebook themselves and were wondering if it could be integrated into the use case. It really got me thinking. Maybe reMarkable or something like it could help bridge a student's writing with an LLM without having to fall back t…

I've written a bunch of tools around the remarkable, and it can be easily integrated here, although realtime integration is quite fussy. At this point, it's easiest to vibe code your own, if you're interested, start with https://github.com/ddvk/rmapi and look into pulling documents and optionally putting them back.

remarkable keeps inking data on a per-page basis in vector format, so there's a bit of work to be done to render the inking and then visually assessing it, but it's all pretty much "done" work.

Remarkable's cloud syncing tech is .. not great, and I'd speculate you'd get a better setup here with some sort of harness where you talk / text with an agent that has access to all the files, dates, OCRs, etc first before you consider trying to generate PDFs / pages and put them back onto the tablet.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#115
post #103
post #58

Earlier quoted context omitted.

Worse, because students complained about the difficulty of the AI-graded quizzes, they switch to multiple-choice questions only, which increases engagement, but after analyzing the exam results they determine that multiple-choice questions don't seem to help and add AI-graded questions back, after which engagement drops again. That means their experiment design is partially caused by their results instead of the othe…

I agree with all the criticism, but I'd like to point out that this kind of study must be done in a way that doesn't discriminate any of the students. It might even be considered unethical to withhold a tool that would be already available just to do research, once there is at least some evidence or very strong suspicion that it might be, in fact, beneficial for at least some of the learners (i.e., you cannot forbid…

> but I'd like to point out that this kind of study must be done in a way that doesn't discriminate any of the students.

This is definitionally impossible. Testing whether the tool benefits the students that get it is an attempt to benefit certain students at the expense of others. Get over this hangup if you want to do this sort of research.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#116

Earlier quoted context omitted.

I used to TA a graduate level CS math class at Georgia Tech. We regularly saw that the students who self-organized study groups did dramatically better in the course than average. One semester they told us to put everyone in study groups to see if it helped. The effect disappeared. Turns out that it was the self-selection of the most engaged students into a small group that mattered, not the study group itself.

So there might be zero effect? If it's purely a correlation, then maybe those students would be more successful than average even without the study group. They're already the most motivated kids. Maybe they just do "motivated kid stuff" and would still outperform.

Yes, that’s what I think at this point. There is no effect of the study group except as a support group. (That’s all it was for me when I was a student and joined the self-organized study group.)

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#117
post #71

Earlier quoted context omitted.

> Nobody creates cards on their phone, or while they're walking Wait, when are you doing it then? No wonder you think it sucks! Adopt some modern tools, yo. Use Anki, or vibe code your own app.

> Use Anki Anki is great for studying, but the card creation experience sucks. To be specific: I found creating any custom card type immediately dropped me into the bowels of CSS. It felt like writing HTML by hand. Is there any facility for re-using shared pieces? I felt like it needed a static site generator type tool to move up a layer of abstraction and reduce the copying of chunks into my card. Is there one? Plea…

I rarely make cards by hand anymore. I would recommend forking something like https://github.com/jasperket/clanki and editing it (perhaps with an agent) so that it works exactly to your liking.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#118
post #60

A lot of pessimism in the comments, but I am just happy that we are seeing some work towards bridging the 2 Sigma gap for regular education vs. elite private tutoring. I can't imagine that people assume it's the physical presence of the tutor that is making the difference, it has to come down to the personalisation and expertise which is exactly what AI can provide in a form. And yea it might not be "there" yet. But…

Tell me if I am oversimplifying, but I never understood the noise about the two sigma problem. Like, of course if you have a private tutor to immediately answer any question that pops into your head at the immediate moment you get confused, you are going to learn vastly more efficiently than in a large classroom where once you get confused you are likely to stay confused. To say nothing of how the pace will likely ei…

The noice isn’t about the effect, it’s about trying to scale the outcome. We essentially have proof of a method that is significantly better than what is used with the broad population, but we cannot scale it to the broad population. That’s the “problem” in the two sigma problem.

And in general ongoing from “it’s obvious to me” to a quantized effect is always an effort. For instance I don’t believe you are correct in assuming that one would see a two sigma effect in physical exercise when comparing someone who attends a regular exercise course va someone who spends the same time with a personal trainer. Two sigma is a lot, and you won’t be lifting significantly heavier weights from hav a PT vs doing starting strength training. In my opinion you would most definitely see a two sigma effect from doing steroids though. But this is all pure speculation which underlines that part of the “noice” is about the documented aspect of the two sigma problem, where we have a body of data to work against not just personal assumptions. That said if anyone has studies on personal trainers va course work in fitness I’d love to see my world image challenged by data.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#119
post #36
post #29

Yes! Very exciting to see this. Bloom's Two Sigma Opportunity suggests that there's another SD improvement available: https://en.wikipedia.org/wiki/Bloom%27s_2_sigma_problem

The story around Bloom's two sigma is a bit complex https://nintil.com/bloom-sigma/

Apparently my definition of tutor is so demanding that it should not be expected at all from a tutor (even if provided with information on the tutees weaknesses, tutors don't use the information effectively). However, the point of these studies are to gauge large scale effects rather than high variance small sample size effects. There's a lot more that is very interesting.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#120
post #70
post #38

Earlier quoted context omitted.

I work in consulting and one of my projects is piloting an AI use case for a department within one of my clients. On a discovery call someone casually brought up that they bought a reMarkable notebook themselves and were wondering if it could be integrated into the use case. It really got me thinking. Maybe reMarkable or something like it could help bridge a student's writing with an LLM without having to fall back t…

> Maybe reMarkable or something like it could help bridge a student's writing with an LLM without having to fall back to a laptop or ipad What does “bridge a student’s writing” mean?? If this is a real argument it needs to be clearer. What’s the functional difference between a Remarkable and an iPad? The former is less responsive, costs less, and has better battery life, right? I really don’t see how that’s significa…

> costs less

I assumed that too, back when i thought that getting an e-ink pad would be cool.

Right now Amazon is quoting me $450 on a 10.3" reMarkable, but the price goes up for the 11.8" model and/or for bells and whistles. It looks like an 11" ipad is $450 and up (based on https://www.apple.com/ipad/compare/), although you need to buy the $80-100 iPencil separately.

So yeah - I was hoping that e-ink things would cost less, but they're still pretty expensive and certainly competitive with the iPad.

Post reply on HN