Live data from Hacker News

Fighting Fire with Fire: Scalable Oral Exams

behind-the-enemy-lines.com

91–100 of 288 posts

Re: Fighting Fire with Fire: Scalable Oral Exams

#91
post #42

I have a lot of complicated feelings and thoughts about this, but one thing that immediately jumps to my mind: was the IRB (Institutional Review Board) consulted on this experiment? If so, I would love to know more details about the protocol used. If not, then yikes!

Turns out that under the USA Code of Federal Regulations, there's a pretty big exemption to IRB for research on pedagogy:

CFR 46.104 (Exempt Research):

46.104.d.1 "Research, conducted in established or commonly accepted educational settings, that specifically involves normal educational practices that are not likely to adversely impact students' opportunity to learn required educational content or the assessment of educators who provide instruction. This includes most research on regular and special education instructional strategies, and research on the effectiveness of or the comparison among instructional techniques, curricula, or classroom management methods."

https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-...

So while this may have been a dick move by the instructors, it was probably legal.

Re: Fighting Fire with Fire: Scalable Oral Exams

#92

Earlier quoted context omitted.

The issue is that it is not scalable, unless there is some dependable, automated way to convert handwriting to text.

Why is this a problem now, but was not a problem for the past few centuries? This class had 36 students, you could grade that in a single evening.

I agree with you and the other posters actually, but I think the efficiency compared with typed work is the reason it’s having such a slow adoption. Another thing to remember is that there is always a mild Jevons paradox at play; while it's true that it was possible in previous centuries, teacher expectations have also increased which strains the amount of time they would have grading handwritten work.

Re: Fighting Fire with Fire: Scalable Oral Exams

#93
post #4

Being interrogated by an AI voice app... I am so grateful I went to university in the before time If this is the only way to keep the existing approach working, it feels like the only real solution for education is something radically different, perhaps without assessment at all

As others have pointed out the radical new approach will simply be reverting to the approach before networked computing took off. Hand written exams at a set time and placed graded by hand by human graders.

Re: Fighting Fire with Fire: Scalable Oral Exams

#94
post #76

Earlier quoted context omitted.

The issue is that it is not scalable, unless there is some dependable, automated way to convert handwriting to text.

University exams being marked by hand, by someone experienced enough to work outside a rigid marking scheme, has been the standard for hundreds of years and has proven scalable enough. If there are so many students that academics can’t keep up, there are likely too many students to maintain a high standard of education anyway.

> there are likely too many students to maintain a high standard of education anyway.

Right on point. I find particularly striking how little is said about whether the best students achieve the best grades. Authors are even candid that different LLMs asses differently, but seem to conclude that LLMs converging after a few rounds of cross reviews indicate they are plausible so who cares. The apparences are safe.

Re: Fighting Fire with Fire: Scalable Oral Exams

#95
post #45

Humanization and responsibility issues aside (I worry that the author seems to validate AIs judgement with no second thought) education is one sector which isn't talked about enough in terms of possible progress with AI. Ask about any teacher, scalability is a serious issue. Students being in classes above and under their level is a serious issue. non-interactive learning, leading to rote memorization, as a result of…

I don’t understand.

Isn’t the poor performance on those exercises also part of their overall performance? Do you mean just that their positive work outweighs the bad work?

Re: Fighting Fire with Fire: Scalable Oral Exams

#96

I predict by the very next semester students still be weaponizing Reasonable Accommodation requests against any further attempts at this

Universities are rapidly becoming useless as a signal of knowledge and competency of their graduates.

Re: Fighting Fire with Fire: Scalable Oral Exams

#97
post #2

If you can use AI agents to give exams, what is stopping you from using them to teach the whole course? Also, with all the progress in video gen, what does recording the webcam really do?

What's stopping you from just using the AI to directly accomplish the ultimate goal, rather than taking the very indirect route of educating humans to do it?

Yes I feel like we still don’t have a good explanation for why AI is super human at stand alone assessments but fall down when asked to perform long term tasks.

Re: Fighting Fire with Fire: Scalable Oral Exams

#98

This is all so crazy to me. I went to school long before LLMs were even a Google Engineer's brianfart for the transformer paper and the way I took exams was already AI proof. Everything hand written in pen in a proctored gymnasium. No open books. No computers or smart phones, especially ones connected to the internet. Just a department sanctioned calculator for math classes. I wrote assembly and C++ code by hand, and…

TFA's case involved examinations about the student's submitted project work. It's not the same thing. Even for a more traditional examination with no such context attached one might still want to rely on AI for grading. (Yeah, I know, that comes across as "the students are not allowed to use AI for cheating, but the profs are!".) Also, IMO oral examinations are quite powerful for detecting who is prepared and who isn…

> On the down side they also help the extroverts and the confident, and you have to be careful about preventing a bias towards those.

This is true, but it is also why it is important to get an actual expert to proctor the exam. Having confidence is good and should be a plus, but if you are confident about a point that the examiner knows is completely incorrect, you may possibly put yourself in an inescapable hole, as it will be very difficult to ascertain that you actually know the other parts you were confident (much less unconfident) in.

Re: Fighting Fire with Fire: Scalable Oral Exams

#99
This seems like a mistake. On the one hand, other commenters' experiences provide additional evidence that oral communication is a vastly different skill from the written word and ought to be emphasized more in education. Even if a student truly understands a concept, they might struggle at talking about it in a realtime context. For many real-world cases, this is unacceptable. Therefore the skill needs to be taught.

On the other hand, can an AI exam really simulate the conditions necessary for improving at this skill? I think this is unlikely. The students' responses indicate not a general lack of expertise in oral communication but also a discomfort with this particular environment. While the author is making steps to improve the environment, I think it is fundamentally too different from actual human-to-human discussion to test a student's ability in oral communication. Even if a student could learn to succeed in this environment, it won't produce much improvement in their real world ability.

But maybe that's not the goal, and it's simply to test understanding. Well, as other commenters have stated, this seems trivially cheatable. So it neither succeeds at improving one's ability in oral communication nor at testing understanding. Other solutions have to be thought of.

Re: Fighting Fire with Fire: Scalable Oral Exams

#100

Just let students use whatever tool they want and make them compete for top grades. Distribution curving is already normal in education. If an AI answer is the grading floor, whatever they add will be visible signal. People who just copy and paste a lame prompt will rank at the bottom and fail without any cheating gymnastics. Plus this is more like how people work. https://sibylline.dev/articles/2025-12-31-how-agent-…

I think the real problem is that AIs have super human performance on one off assessments like exams, but fall over when given longer term open ended tasks.

This is why we need to continue to educate humans for now and assess their knowledge without use of AI tools.

Post reply on HN