I'm on record saying that a system like this with some extra hardware (i.e. a way for the LLM to have live understanding of the student's paper notebook or handout which are being written in with a plain old pencil) combines the best of both worlds - individual tutoring with approximately zero screen time which scales linearly with the number of students. The role of the teacher or professor then becomes a manager of…
A 'smart pen' that records the student's writing in some way, maybe? My first thought was a tablet that boots straight into a writing software but students should not be subjected to any amount of latency in their writing. Practically, I think if you want the AI system to have a live view of what the student's doing you're going to have to replace one of either the tablet or the writing instrument. A wearable camera…
New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
31–40 of 126 posts
Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#32Shocking that a well executed AI tutor improves outcomes. Hasn't computer assisted interactive learning already been proven for years? Why does there seem to be so much skepticism about enhancing it with AI? Is this just something like, astoundingly slow adoption or poor execution? Being held back by paper textbook makers? Teachers unions dragging their feet? How can interactive AI driven individually paced learning…
Motivation is also a huge part of the problem. I'm wondering if the novelty of the AI tutoring gets more people to try it and whether it would wear off?
It's surprising to me that many students at Dartmouth don't read the textbook. You'd think college admissions would select for that?
It seems promising but, as they say, more research needed.
Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#33Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#34Interesting, congrats. Are you planning on opening access to Phosphor?
Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#35Earlier quoted context omitted.
A 'smart pen' that records the student's writing in some way, maybe? My first thought was a tablet that boots straight into a writing software but students should not be subjected to any amount of latency in their writing. Practically, I think if you want the AI system to have a live view of what the student's doing you're going to have to replace one of either the tablet or the writing instrument. A wearable camera…
there was a pen that used special paper to directly record your notes (15-20years ago)... should be possible nowadays to directly transfer this to a connected device and have it feed it to an llm. and after looking it up, it appears they are still available: https://www.livescribe.com/landingpage/ls3_onenote/
Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#36Yes! Very exciting to see this. Bloom's Two Sigma Opportunity suggests that there's another SD improvement available: https://en.wikipedia.org/wiki/Bloom%27s_2_sigma_problem
Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#37Shocking that a well executed AI tutor improves outcomes. Hasn't computer assisted interactive learning already been proven for years? Why does there seem to be so much skepticism about enhancing it with AI? Is this just something like, astoundingly slow adoption or poor execution? Being held back by paper textbook makers? Teachers unions dragging their feet? How can interactive AI driven individually paced learning…
There ARE technologies that have improved things, but so much high-cost useless tech has been shoved into every level of education that many educators are incredibly leery of new tech.
The issue is that while the underlying technology is useful, the way it gets integrated is frequently not. An administrator cuts a deal for a product they never have to use to an ed-tech giant for a huge amount. Because the ink is dry and a huge sum of money has been spent admins pressure educators to use the technology as much as possible regardless of outcome.
In that context it makes a lot more sense why there is pushback and FUD among educators.
Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#38I'm on record saying that a system like this with some extra hardware (i.e. a way for the LLM to have live understanding of the student's paper notebook or handout which are being written in with a plain old pencil) combines the best of both worlds - individual tutoring with approximately zero screen time which scales linearly with the number of students. The role of the teacher or professor then becomes a manager of…
Maybe reMarkable or something like it could help bridge a student's writing with an LLM without having to fall back to a laptop or ipad.
Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#39Conflicted about this study. On one hand, LLMs have been incredible for my personal learnings of new concepts. On the other, I'm sceptical of that it'll have "strong benefits" at scale; I'd be more in favor if the wording was "some"/"moderate". I reckon self-selection plays a huge part, as mentioned in the "Limitations" section of the paper. I'd also caution against attaching the tool to grading. That means students…
Mind if I ask what did you learn and how you're using it?
The reason I'm asking is that I repeatedly felt excitement only to realize down the line that the explanations didn't actually translate into practical skills. I'm not sure it's even an AI problem, it's a "doing versus reading" problem. Same as with reading a pop-science article and thinking to myself that I learned something about physics or medicine or mathematics.
Re: New AI tutor achieves 0.71-1.30 SD effect size in Dartmouth course [pdf]
#40First, the headline result of 0.7*sigma improvement is the output of a statistical based on lessons/reviews they engaged with and their mid-term score, with that shift being for "full engagement". Based on their tables something like ~16 students (11% of the group) actually reached that level of engagement
Second, trying to incorporate past grades into their modelling is not a substitute for a randomized trial.
Third, the headline engagement number of 90% is for "engaging with the platform, via Module Review or Lesson Quizzes, at least once". I don't know why much of that couldn't just be attributed to novelty. Or even partly a professor with all sorts of enthusiasm for the platform.
Fourth, the "full dosage" effectiveness is measured based the final exam scores. Were these exam questions produced independently from the "Phosphor" materials? (e.g. by blinding?) Were they checked for direct overlap with those materials? The 0.7 sigma shift is 3 points on a 24 point exam; if even a few of the questions on that exam were very similar to those materials it could account for almost all of it. This is not clear to me from the manuscript.
If this was the case, then it's a question less of "is AI effective" vs. "did the students look at the materials". You could still argue that the AI platform got them to read, but that is a somewhat different statement than the AI helped them learn.