Live data from Hacker News

Ontario auditors find doctors' AI note takers routinely blow basic facts

theregister.com

51–60 of 141 posts

Re: Ontario auditors find doctors' AI note takers routinely blow basic facts

#51
post #41

Earlier quoted context omitted.

They do know what they don't know. There's a probability distribution for outputs that they are sampling from. That just isn't being used for that purpose.

Oh, you mean somewhere it is tracking the statistical likelihood of the output. Yeah I buy that, although I think it just tends towards the most likely output given the context that it is dragging along. I mean it wouldn’t deliberately choose something really statistically unlikely, that’s like a non sequitur.

From its point of view what does it mean "to know".

Is it the token (or set of tokens) that are strictly > 50% probable or is it just the highest probability in a set of probabilities?

While generating bullshit is not ideal for a lot of use cases you don't want your premier chat bot to say "I don't know" to the general public half the time. The investment in these things requires wide adoption so they are always going to favour the "guesses".

Re: Ontario auditors find doctors' AI note takers routinely blow basic facts

#52

Earlier quoted context omitted.

That's just not how they work, really. They don't know what they don't know and their process requires an output. I think they're getting better at it, but it's likely just the number of parameters getting bigger and bigger in the SOTA models more than anything.

They do know what they don't know. There's a probability distribution for outputs that they are sampling from. That just isn't being used for that purpose.

I’m not clear what you mean by “know.” If you mean “the information is in the model” then I mostly agree, distributional information is represented somewhere. But if you mean that a model can actually access this information in a meaningful and accurate way—say, to state its confidence level—I don’t think that’s true. There is a stochastic process sampling from those distributions, but can the process introspect? That would be a very surprising capability.

Re: Ontario auditors find doctors' AI note takers routinely blow basic facts

#53

Anecdotally, we use an LLM note-taker at work for meetings. I had to intervene recently because our CIO was VERY angry at our vendor for something they promised to do and never did. He wasn't at the meeting where the "promise" was made. I was. They never promised anything, and the discussion was significantly more nuanced than what the LLM wrote in the detailed summary. In other cases, I have seen it miss the mark wh…

> I would think for compliance reasons hospitals would not want to alter the records and only go by transcripts, but what do I know... I'm puzzled by this as well. Why not just generate a transcript and be done with it? If it's a particularly long transcript that's being referenced repeatedly for whatever reason let the humans manually mark it up with a side by side summary when and where they feel the need. At least…

I mean the reasons are the same AI is being pushed everywhere.

The businesses offering these services want to say "we are using AI" to their stake holders and the government committees who approve this shit don't have the skills or knowledge to evaluate the effectiveness in addition to the fact they likely don't even use the tools they have approved for use.

Re: Ontario auditors find doctors' AI note takers routinely blow basic facts

#54

Earlier quoted context omitted.

When designing AI-based user experiences I refer to this as provenance. It’s a vital aspect of trust, reliability, compliance and more. If a software system includes LLM output like this but doesn’t surface the provenance of its output for human evaluation and verification then it’s at best poor user experience, and at worst a dangerous one.

At the same time, do you really want every conversation you have with your doctor recorded, handed over to third party companies, and stored forever with your medical file? Plus what doctor has time to sit down and re-listen to your visit to check to make sure the AI didn't screw up at some point in the future anyway? If your doctor isn't going to be verifying the accuracy from those recordings who would? Overseas co…

>At the same time, do you really want every conversation you have with your doctor recorded

Yes. This is what medical records are. They've been kept by doctors for a reason.

It's not like the doctor is talking to you about which anime series are the best. You're talking about your health, your body, your disease, your treatment.

It's important to keep track of that.

>Plus what doctor has time to sit down and re-listen to your visit to check to make sure the AI didn't screw up at some point in the future anyway

No doctor.

Which is why it really should be their (or their assistant's) job to record the relevant parts of the conversations.

>At what point does it become a larger waste of time and money to babysit an incompetent AI than just not using one in the first place?

At this point, as the audit shows.

Except the industry (both the AI vendors and healthcare) are going YOLO¹ and relying on AI anyway.

>There are some good uses for AI, but I'm not convinced that this (or many other cases where accuracy matters) is one of them.

This has always been the case, but the marketing now has reached of point of gaslighting in trying to make people collectively forget that or pretend that it's not the case.

Once hard evidence is presented (like in this case), the defense is invariably that it's a temporary quality issue that's going to be resolved as the AI improves Any Day Now™, and that it's wise to live as if it were the case already² (and everyone who disagrees is a fool that Will Be Left Behind™).

The level of fervor in this rhetoric gives me an impression that the flaw is so fundamental that it won't be fixed in any form of AI based on today's technologies, that the AI vendor leadership knows this, and that the entire industry is, at this point, is a grand pump-and-dump scheme.

I hope I'm wrong.

____

¹ See, you only live once. But there are millions of you. So, like, whatever if you don't. Something something economies of scale to them.

² This is called a phantasm.

Re: Ontario auditors find doctors' AI note takers routinely blow basic facts

#55
post #52

Earlier quoted context omitted.

They do know what they don't know. There's a probability distribution for outputs that they are sampling from. That just isn't being used for that purpose.

I’m not clear what you mean by “know.” If you mean “the information is in the model” then I mostly agree, distributional information is represented somewhere. But if you mean that a model can actually access this information in a meaningful and accurate way—say, to state its confidence level—I don’t think that’s true. There is a stochastic process sampling from those distributions, but can the process introspect? Tha…

yes:

> In this experiment, however, the model recognizes the injection before even mentioning the concept, indicating that its recognition took place internally.

https://www.anthropic.com/research/introspection

Re: Ontario auditors find doctors' AI note takers routinely blow basic facts

#56
Yep. It happened to me just recently.

Diagnosed with Runner's Knee.

AI summary said I was diagnosed with osteoporosis, and had hip pain and walking difficulty, though literally none of that was ever said or implied.

CHECK YOUR TRANSCRIPTS. Always, but especially with LLM transcribers, which fairly frequently include common symptoms which don't exist, or claim a diagnosis which is common and fits a few details but not others. Get them fixed, it can very strongly affect your care and costs later if it's wrong.

Anecdotally, I'd say that outside of a couple very simple and very common things, about 50% of the "AI" summaries I've had have been wrong somewhere. Usually claiming I have symptoms that don't exist, occasionally much more serious and major fabrications like this time.

LLMs are NOT normal speech to text software, and they shouldn't be treated like one. They'll often insert entire sentences that never occurred. In some contexts that might be fine, but definitely not in medical records.

Re: Ontario auditors find doctors' AI note takers routinely blow basic facts

#57

Anecdotally, we use an LLM note-taker at work for meetings. I had to intervene recently because our CIO was VERY angry at our vendor for something they promised to do and never did. He wasn't at the meeting where the "promise" was made. I was. They never promised anything, and the discussion was significantly more nuanced than what the LLM wrote in the detailed summary. In other cases, I have seen it miss the mark wh…

Every doctor's visit I've had, I have been able to make corrections to the record afterward, because there have been meaningful mistakes almost half the time.

ALWAYS check your summaries immediately, and contact your doctor ASAP. They can generally fix it themselves, and it's best done when everyone still has some memory of the event.

Re: Ontario auditors find doctors' AI note takers routinely blow basic facts

#58

Earlier quoted context omitted.

That's just not how they work, really. They don't know what they don't know and their process requires an output. I think they're getting better at it, but it's likely just the number of parameters getting bigger and bigger in the SOTA models more than anything.

They do know what they don't know. There's a probability distribution for outputs that they are sampling from. That just isn't being used for that purpose.

Well, with thinking models, it’s not that simple. The probability distribution is next token. But if a model thinks to produce an answer, you can have a high confidence next token even if MCMC sampling the model’s thinking chain would reveal that the real probability distribution had low confidence.

Re: Ontario auditors find doctors' AI note takers routinely blow basic facts

#59
post #56

Yep. It happened to me just recently. Diagnosed with Runner's Knee. AI summary said I was diagnosed with osteoporosis, and had hip pain and walking difficulty, though literally none of that was ever said or implied. CHECK YOUR TRANSCRIPTS. Always, but especially with LLM transcribers, which fairly frequently include common symptoms which don't exist, or claim a diagnosis which is common and fits a few details but not…

I've actually seen this lead to serious issues when a zoom LLM summary attributed statements to someone who didn't say them.

Someone else who couldn't attend the meeting later read that summary and it created a major argument because the topic had been a sore subject for this person due to an ongoing debate at the company. Everyone who attended the meeting confirmed it was an error, but the coincidental timing made it hard for him to accept, because the LLMs summary presented things in a way that validated this person's concerns that had been previously minimized by some folks on that meeting.

The drama got heated to the point where management produced a policy about not trusting generative output without independent verification. Seems at least it was a lesson learned.

Re: Ontario auditors find doctors' AI note takers routinely blow basic facts

#60
post #41

Earlier quoted context omitted.

They do know what they don't know. There's a probability distribution for outputs that they are sampling from. That just isn't being used for that purpose.

Oh, you mean somewhere it is tracking the statistical likelihood of the output. Yeah I buy that, although I think it just tends towards the most likely output given the context that it is dragging along. I mean it wouldn’t deliberately choose something really statistically unlikely, that’s like a non sequitur.

Well, it's not tracking. As it predicts each token it is sampling from a probability distribution -- that's what the matrix multiplies are for. It gets a distribution over all tokens and then picks randomly according to that distribution. How flat or how spiky that distribution is tells you how confident it is in its answer.

But it then throws that distribution away / consumes it in the next token calculation. So it's not really tracking it per se.

Post reply on HN