Live data from Hacker News

New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

intextbooks.science.uu.nl

91–100 of 126 posts

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#91
post #71

Earlier quoted context omitted.

> Nobody creates cards on their phone, or while they're walking Wait, when are you doing it then? No wonder you think it sucks! Adopt some modern tools, yo. Use Anki, or vibe code your own app.

> Use Anki Anki is great for studying, but the card creation experience sucks. To be specific: I found creating any custom card type immediately dropped me into the bowels of CSS. It felt like writing HTML by hand. Is there any facility for re-using shared pieces? I felt like it needed a static site generator type tool to move up a layer of abstraction and reduce the copying of chunks into my card. Is there one? Plea…

Write a tool to reduce the friction.

I use an Emacs based SRS tool. I have a capture template to quickly make a card.

For anything tedious, it's critical to reduce the friction!

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#92

I am somewhat skeptical of this. First, the headline result of 0.7*sigma improvement is the output of a statistical based on lessons/reviews they engaged with and their mid-term score, with that shift being for "full engagement". Based on their tables something like ~16 students (11% of the group) actually reached that level of engagement Second, trying to incorporate past grades into their modelling is not a substit…

It feels to me that the venn diagram between "students that fully engaged with the material" and "students that learned well from the material" is going to basically be a circle for any teaching method.

The tricky question at the forefront of education research is, at least in my mind, trying to thread the needle between effective techniques that students don’t like, ineffective techniques that students do like, and poorly defined techniques that administrators like. And on top of all that student self-reports aren’t actually very reliable indicators of learning progress at all!

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#93
post #71

Earlier quoted context omitted.

> Nobody creates cards on their phone, or while they're walking Wait, when are you doing it then? No wonder you think it sucks! Adopt some modern tools, yo. Use Anki, or vibe code your own app.

> Use Anki Anki is great for studying, but the card creation experience sucks. To be specific: I found creating any custom card type immediately dropped me into the bowels of CSS. It felt like writing HTML by hand. Is there any facility for re-using shared pieces? I felt like it needed a static site generator type tool to move up a layer of abstraction and reduce the copying of chunks into my card. Is there one? Plea…

I really never thought it would come up in a HN thread but I’m actually working on a modified version of Anki as a personal project (not quite ready yet though, but will open source probably in a few weeks) where improving the editing/curating/creating experience is a big focus. I’m trying to make it markdown-based too.

Just to pick your brain real fast, when creating new cards (from scratch?) what would a better experience look like for you?

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#94
post #36
post #29

Yes! Very exciting to see this. Bloom's Two Sigma Opportunity suggests that there's another SD improvement available: https://en.wikipedia.org/wiki/Bloom%27s_2_sigma_problem

The story around Bloom's two sigma is a bit complex https://nintil.com/bloom-sigma/

Beat me to it. The Wikipedia page for example looks like an advertisement for tutoring firms, and when you dig into the sources, it’s pretty oversold

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#95

I am somewhat skeptical of this. First, the headline result of 0.7*sigma improvement is the output of a statistical based on lessons/reviews they engaged with and their mid-term score, with that shift being for "full engagement". Based on their tables something like ~16 students (11% of the group) actually reached that level of engagement Second, trying to incorporate past grades into their modelling is not a substit…

Education research is hard hard hard. Getting clean studies, especially at large sample sizes, is extremely difficult, leading to a lot of ambiguity in results.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#96
post #11

I'm on record saying that a system like this with some extra hardware (i.e. a way for the LLM to have live understanding of the student's paper notebook or handout which are being written in with a plain old pencil) combines the best of both worlds - individual tutoring with approximately zero screen time which scales linearly with the number of students. The role of the teacher or professor then becomes a manager of…

Today I saw a demo of Remarkable turned into Voldemort's diary from Harry Potter - you write to it, and it writes back, in handwriting.

This has been around for a bit, not sure it writes back in handwriting.

https://github.com/awwaiid/ghostwriter

Installation not for the average user.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#98

I am somewhat skeptical of this. First, the headline result of 0.7*sigma improvement is the output of a statistical based on lessons/reviews they engaged with and their mid-term score, with that shift being for "full engagement". Based on their tables something like ~16 students (11% of the group) actually reached that level of engagement Second, trying to incorporate past grades into their modelling is not a substit…

Thank you for the feedback! Maybe the following info will be helpful when considering our results:

1. Quiz completion is our deliberately conservative lower bound on reading compliance, and the 0.71 figure is not a claim that those 16 students each gained that much. The estimate is from a regression carried by the per-lesson slope, fit across the whole dosage distribution, and the underlying dosage-performance relationship is essentially unchanged whether or not zero-completion students are included (R² 0.091 vs 0.096). In other words, more Phosphor use is strongly associated with better performance across the whole range of usage - not just the group who completed all content.

The numbers in Table 1 show how dosage was distributed across the course. We report that across the class, the median percentage of lessons reached on Phosphor, including both students with an account and those who never logged in, was above 90%. Among platform users in particular, it was 96%.

We'd like to emphasize that for this pilot, the platform was presented to students as an entirely optional "study aid", and our adoption rates far exceed those reported in the past for optional interventions. It will be interesting to see how things go when we attach completion to the course grade, as we're thinking of doing in the fall. Past literature from interventions in college courses predicts that this will achieve far higher levels of engagement, bringing the high-dosage effect to a large proportion of the class.

2. We explicitly note this in Limitations; it's an observational study. We were unable to do an RCT for this course since it raised an ethics consideration - neither we nor the instructors wanted to deprive students entirely of a course material that could have been helpful for them. We'd love to run a randomized trial at some point though - one way to do this is a crossover, where we offer the treatment to one of two groups, then switch it over to the other midway through the trial, so that both get even treatment. Another possibility is randomly selecting students to get access to MCQ-only vs. CRQ-enabled quizzes. That being said, this mechanism of conditioning on past performance is well-known and relatively robust for observational studies of educational interventions.

3. The platform was created independently of the instructors of the course. The instructors designed their curriculum ahead of time (as had been taught for years of past offerings of the course), lectured in a conventional style, referenced the course's official textbook (Freedman, Pisani, Purves) and suggested homework problems from the textbook only. The instructional content was authored using material that every student in the course had access to, and did not feature exam questions that students were evaluated on after-the-fact.

Phosphor was not endorsed publicly by the instructors, and was spread primarily via student word-of-mouth. In fact, one instructor in this course initially believed the project would be "a waste of time" and refused to collaborate with us for the pilot. Despite this, 97% of the students in this instructor's section used the platform!

Engagement also persisted well across the full ten-week term, and two-thirds of Review attempts involved retries spaced a day or more apart — not a pattern typically produced by novelty effects.

4. Instructors wrote exams independently with their long-running FPP-based curriculum. Even if we steelman and suppose that "the platform just got students to engage" rather than truly learn, this is refuted by our result that the MCQ-only Module 2 had similar engagement but no dosage relationship. This strongly suggests that the CRQ format was a driver of the results.

As we mention in the paper, we agree that replication, especially across contexts, is a priority. For us in particular, this means not only across other courses, but across other institutions as well. And an RCT would certainly help lock in the causal claim.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#99
post #87
post #82

Earlier quoted context omitted.

Hmm, it might be better to just go straight to what leading countries currently do (Finland/Estonia/Japan/Singapore/etc). Other factors probably play a bigger role, like school funding, school autonomy, teacher professionalism (high education level + continuing education + good pay), free school meals/healthcare/transportation etc

One major limitation is the cost. It costs lots of money to train teachers and pay them good salaries. What leading countries are doing currently might not be sustainable for the long run. Cognitive abilities also different greatly from people to people in different discipline. There is no silver bullet.

The long term sustainability in at least one of those leading countries (Japan) has been shown effective for at least the period from the end of WWII to now. The principle difficulty they have is not maintaining high levels of educational success but in having enough children to keep the schools open and the teachers fully employed. But, that is a different concern than successful educational outcomes.

The average level of capability and comprehension in fundamental disciplines for students completing primary is a direct result of some fundamental differences in the way they approach classroom organization: for one middle school and high school students do not change classrooms during a school day.

Teachers rotate rooms while students remain in place and this maintains a less disjointed, less distracted transition between classes. There are still very little to zero technological advances to the teaching methods: blackboards/whiteboards, overhead projectors, hand written paperwork, textbooks.

The reasoning is simple. The fundamentals which students in primary education are in need of learning do not change much. Last years' textbooks are perfectly useful and the cost of replacement is directly carried by each student. If they damage a book, they must replace it.

Salaries for teachers are not unusually high, but they are also not low. The building and administrative costs are kept low, for one, by there not being a significant non-instructor labor expense of janitorial and maintenance workers. Schools share the pool of municipal HVAC and other trades for serious infrastructure, but janitors? Nope. The students are required to actually clean up after themselves, and to actually clean the whole school. It makes a difference, and those avoidable labor costs can be directed to proper compensation for the teaching staff.

Anecdotal observation of the overall efficacy of this approach reaches areas not usually measured as 'educational success' but also includes one noteworthy artifact. The common clerk, shop keeper, cook, or gas tank delivery driver all know that i = sqrt(-1) and that a complex number is a pair of numbers, one of which is the coefficient of (i). Let that sink in. How many people graduate with a B.A. from western schools without that?

There is not a 'silver bullet' but there are a full compliment of approaches which when applied together, consistently, and persistently, yield excellent sustainable outcomes.

Back to the falling population problem. There are pluses and minuses from the choice to provide separate boys and girls middle and high schools; one of which is lower teenage pregnancy and it's inverse, lower adult birthrates. (It's not the only reason for the lower birthrates, but it's not having zero impact.) You cant win 'em all.

Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]

#100
post #42

The title is misleading. This isn't an AI tutor so much as a practice quiz platform with an AI autograder. > constructed-response questions (CRQ) are graded by Claude Sonnet 4.6 against instructor-defined, question-specific rubric criteria > Crucially, LLMs make it feasible to grade formative CRQ against rubric criteria at scale, a capability that appears pedagogically significant rather than merely convenient. They…

> a practice quiz platform with an AI autograder. What do you think tutoring is?

The role of a tutor is to find your weakest point (or points) and give you personalized advice on how to improve on them. It is important that you can put trust on the tutor, as your weak point is likely due to a blind spot. When given criticism, it can be viewed as objective or subjective, i.e. a "question of taste", and thus the criticism would be viewed as invalid and not acted upon. As to whether the tutor should point your weakest point or some combination of weak point, it should depend on what helps you learn the most efficiently; it might be to focus on one aspect, or work in a more integrated fashion. Last (although that may be first), they should consider what are really your learning objectives to tailor that advice.
Post reply on HN