Live data from Hacker News

Don't ask an LLM for a confidence score

justinflick.com

11–20 of 46 posts

Re: Don't ask an LLM for a confidence score

#11

“That’s real, and it’s interesting.” Stop with the slop.

Unfortunately I think that we need to be much more judgmental than is tolerated on this forum if we want to get people to stop writing with AI assistance.

People are not going to stop spamming you with slop text unless the narrative is "Using AI for writing is a mark of poor literacy/low intelligence".

Re: Don't ask an LLM for a confidence score

#12
This resonates with what I'm seeing in healthcare right now and the bad taste it's leaving.

There are a whole host of new EMRs popping up that aim to help clinicians make judgement calls about how to answer certain questions and even when particular procedures are relevant. The last one we demoed, each piece of information it retrieved from the LLM has a confidence score attached to it. Our nurses have to fill out 200+ question forms when taking on a new patient that all HAVE to be completed in a single-go meaning that you can't split it up into multiple forms and it's gotta be one cohesive unit.

Imagine being a nurse with little technical skill and almost no idea how these tools work trying to make sense of what the difference between a 90% and 70% is across 200 different questions. "We're 60% sure the patient is allergic to nuts" means jack shit to them. Granted, sometimes the scores are complimented with actual references in the underlying documentation (history and physical, referring info) but sometimes it's not.

Re: Don't ask an LLM for a confidence score

#13
You don’t ask an LLM, but certain LLMs expose internal metrics you can tell how many token candidates where there what was the score which one and which one was selected.

So there are objective ways to control hallucinations as well as figuring out how “correct” the answer is to some extent.

Re: Don't ask an LLM for a confidence score

#14
I find the recommendation in favour of discrete categories over numeric probability for confidence level intriguing, because it goes against the grain of all advice I've read for humans interested in forecasting.

The easiest argument in favour of numeric probability is that it's simple to evaluate. Someone claiming to be 90 % certain better be right about nine out of ten times -- no fewer, but also not too often! This is useful when you're learning to judge your own confidence, which most people are bad at. (Almost everyone can train themselves to be better at i but some people seem to be good at it with no training.)

The typical argument against categories is that people can mean very different things with the same word: https://i.ibb.co/kQT4Ymz/q.png

It seems like TFA gets around these problems by creating a translation table between a fixed set of categories and the probabilities implied by the model for those categories. That's an interesting approach!

Re: Don't ask an LLM for a confidence score

#15

This resonates with what I'm seeing in healthcare right now and the bad taste it's leaving. There are a whole host of new EMRs popping up that aim to help clinicians make judgement calls about how to answer certain questions and even when particular procedures are relevant. The last one we demoed, each piece of information it retrieved from the LLM has a confidence score attached to it. Our nurses have to fill out 20…

Healthcare still uses fax. No surprise they arent using Frontier AI to answer these questions. (And if you still don't trust AI, have the report on the left side of the screen and the source pages on the right side of the screen)

Re: Don't ask an LLM for a confidence score

#19

Who can read that text color and background combo? I had to turn on reader mode in Firefox.

Firefox on desktop renders as black text on white background. The link colors are yellow, which is not great, but still okay for me to read.

I use the "Dark Background and Light Text" add-on. Sometimes the add-on makes a bad website worse (like IMDB), but it can be switched off with a click.

Re: Don't ask an LLM for a confidence score

#20
post #15

This resonates with what I'm seeing in healthcare right now and the bad taste it's leaving. There are a whole host of new EMRs popping up that aim to help clinicians make judgement calls about how to answer certain questions and even when particular procedures are relevant. The last one we demoed, each piece of information it retrieved from the LLM has a confidence score attached to it. Our nurses have to fill out 20…

Healthcare still uses fax. No surprise they arent using Frontier AI to answer these questions. (And if you still don't trust AI, have the report on the left side of the screen and the source pages on the right side of the screen)

If you work in a field which matters, not trusting even “Frontier AI” is also known as “basic professional competence”. If you’re vibe-coding a marketing site nobody dies if you miss a mistake. The same is not true in healthcare or real engineering, and because AI systems can’t be held accountable you will be.
Post reply on HN