Live data from Hacker News

Fighting Fire with Fire: Scalable Oral Exams

behind-the-enemy-lines.com

211–220 of 288 posts

Re: Fighting Fire with Fire: Scalable Oral Exams

#211

> We surveyed students before releasing grades to capture their experience. [...] Only 13% preferred the AI oral format. 57% wanted traditional written exams. [...] 83% of students found the oral exam framework more stressful than a written exam. [...] > Take-home exams are dead. Reverting to pen-and-paper exams in the classroom feels like a regression. Yeah, not sure the conclusion of the article really matches the…

> they expressed a clear preference for written exams

When I was a student, I would have been quite vocal with my clear preferences for all exams being open-book and/or being able to amend my answers after grading for a revised score.

What I'm saying is, "the students would prefer..." isn't automatically case closed on what's best. Obviously the students would prefer a take-home because you can look up everything you can't recall / didn't show up to class to learn, and yes, because you can trivially cheat with AI (with a light rewrite step to mask the "LLM voice").

But in real life, people really will ask you to explain your decisions and to be able to reason about the problem you're supposedly working on. It seems clear from reading the revised prompts that the intent is to force the agent to be much fairer and easier to deal with than this first attempt was, so I don't think this is a bad idea.

Finally, (this part came from my reading of the student feedback quotes in the article) consider that the current cohort of undergrads is accustomed to communicating mainly via texting. To throw in a further complication, they were around 13-17 when COVID hit, decreasing human contact even more. They may be exceedingly nervous about speaking to anyone who isn't a very close friend. I'm sympathetic to them, but helping them overcome this anxiety with relatively low stakes is probably better than just giving up on them being able to communicate verbally.

Re: Fighting Fire with Fire: Scalable Oral Exams

#212

Earlier quoted context omitted.

If teaching is the goal, a 99% failure rate seems counterproductive.

I'd wager the "Cohorts for programs with a thousand initial students had less than 10 graduates" statement is deceptive, if not outright false. Perhaps lifetimerubyist means "1000 students took the mandatory philosophy and ethics 101 class, but only 10 graduated as philosophy majors"

I believe certain european countries have or had free universities which instead filter students with incredibly difficult courses. Thousands might enter because both tuition and board are free and they would like a degree, but the university ensures that only a small group make it to second year. I believe the filtering is less intense in later years, since the job has already been done by that point.

Re: Fighting Fire with Fire: Scalable Oral Exams

#213
post #166
post #89

Earlier quoted context omitted.

> And why is this a flex exactly? Almost sounds like fraud. Do you think you're just purchasing a diploma? Or do you think you're purchasing the opportunity to gain an education and potential certification that you received said education? It's entirely possible that the University stunk at teaching 99% of it's students (about as equally possible that 99% of the students stunk at learning), but "fraud" is absolute no…

If you have a You could easily raise the bar without sacrificing quality of education (and likely you'd improve it just from the improvement in student:teacher ratio).

Exactly that. Also, I experienced a situation where a free uni (eastern Europe) had low admission criteria and then had a "cleaning" math course, which 80%-90% failed. School still got paid for the number of students admitted, not those who passed.

In another European country, schools get paid for students that passed.

Re: Fighting Fire with Fire: Scalable Oral Exams

#214
post #76

Earlier quoted context omitted.

The issue is that it is not scalable, unless there is some dependable, automated way to convert handwriting to text.

University exams being marked by hand, by someone experienced enough to work outside a rigid marking scheme, has been the standard for hundreds of years and has proven scalable enough. If there are so many students that academics can’t keep up, there are likely too many students to maintain a high standard of education anyway.

The rate of college attendance has increased dramatically in the last 250 years, and especially in the last 75.

In 1789 there were 1,000 enrolled college students total, in a country of 2.8M. In 2025, it is 19M students in a country of 340M. https://educationalpolicy.org/wp-content/uploads/2025/11/251...

In 1950, 5.5% of adults ages 25-34 had completed a 4 year college degree. In 2018, it was 39%. https://www.highereddatastories.com/2019/08/changes-in-educa...

With attendance increasing at this rate (not to mention the exploding costs of tuition), it seems possible that the methods need to change as well.

Re: Fighting Fire with Fire: Scalable Oral Exams

#216
As a University professor, what I really don't get about this "experiment" is the timings. They report:

> 36 students examined over 9 days > 25 minutes average (range: 9–64)

It appears that they examined only 4hrs each day, one student at a time. This is incredibly inefficient.

In my experience, the greatest benefit of doing something like this would be to be able to run these exams in parallel, while retaining a somewhat impartial grading system.

Re: Fighting Fire with Fire: Scalable Oral Exams

#217

Earlier quoted context omitted.

I saw this piece as the start of an experiment, and the use of a "council of AI" as they put it to average out the variability sounds like a decent path to standardization to me (prompt injecting would not be impossible, but getting something past all the steps sounds like a pretty tough challenge) They mention getting 100% agreement between the LLMs on some questions and lower rates on other, so if an exam was compo…

Imagine that LLMs reproduce the biases of their training sets and human data sets are biased against nonstandard speakers with rural accents/dialects/AAVE as less intelligent. Do you imagine their grade won't be slightly biased when the entire "council" is trained on the same stereotypes? Appeals aren't a solution either, because students won't appeal (or possibly even notice) a small bias given the variability of al…

I might be given too much credit, but given the tone of the post they're not trying to apply this to some super precise extremely competitive check.

If the goal is to assess whether a student properly understood the work they submitted or more generally if they assimilated most concepts of a course, the evaluation can have a bar low enough for let's say 90% of the student to easily pass. That would give enough of margin of error to account for small biases or misunderstandings.

I was comparing to mark sheet tests as they're subject to similar issues, like students not properly understanding the wording (and usually the questions and answer have to be worded in pretty twisted ways to properly) or straight checking the wrong lines or boxes.

To me this method, and other largely scalable methods, shouldn't be used for precise evaluations, and the teachers proposing it also seem to be aware of these limitations.

Re: Fighting Fire with Fire: Scalable Oral Exams

#218
It is quite telling, regarding the state of higher education, if actual teachers actually talking to students 1:1 (which is all that an oral exam really needs to be) is brushed away as a non-starter. I can highly empathise with students who feel like the whole enterprise is a farce, and trying to game and cheat that system at every possible turn is the only appropriate response.

Re: Fighting Fire with Fire: Scalable Oral Exams

#219

Earlier quoted context omitted.

I'd wager the "Cohorts for programs with a thousand initial students had less than 10 graduates" statement is deceptive, if not outright false. Perhaps lifetimerubyist means "1000 students took the mandatory philosophy and ethics 101 class, but only 10 graduated as philosophy majors"

I believe certain european countries have or had free universities which instead filter students with incredibly difficult courses. Thousands might enter because both tuition and board are free and they would like a degree, but the university ensures that only a small group make it to second year. I believe the filtering is less intense in later years, since the job has already been done by that point.

Unless you're thinking of huge online courses like Udacity/Coursera, I don't think that's really a thing?

If it is, I'd be fascinated to learn more.

I mean, the logistics would be pretty wild - even a large university's largest lecture theatres might only have 500 seats. And they'd only have one or two that large. It'd be expensive as hell to build a university that could handle multiple subjects each admitting over a thousand students.

Re: Fighting Fire with Fire: Scalable Oral Exams

#220

> We surveyed students before releasing grades to capture their experience. [...] Only 13% preferred the AI oral format. 57% wanted traditional written exams. [...] 83% of students found the oral exam framework more stressful than a written exam. [...] > Take-home exams are dead. Reverting to pen-and-paper exams in the classroom feels like a regression. Yeah, not sure the conclusion of the article really matches the…

That's what so surprising to me - they data clearly shows the experiment had terrible results. And the write up is nothing but the author stating: "glowing success!". And they didn't even bother to test the most important thing. Were the LLM evaluations even accurate! Have graders manually evaluate them and see if the LLMs were even close or were wildly off. This is clearly someone who had a conclusion to promote reg…

> And they didn't even bother to test the most important thing. Were the LLM evaluations even accurate!

This is not true; the professor and the TAs graded every student submission. See this paragraph from the article:

(Just in case you are wondering, I graded all exams myself and I asked the TA to also grade the exams; we mostly agreed with the LLM grades, and I aligned mostly with the softie Gemini. However, when examining the cases when my grades disagreed with the council, I found that the council was more consistent across students and I often thought that the council graded more strictly but more fairly.)

Post reply on HN