Earlier quoted context omitted.
That's just not how they work, really. They don't know what they don't know and their process requires an output. I think they're getting better at it, but it's likely just the number of parameters getting bigger and bigger in the SOTA models more than anything.
They do know what they don't know. There's a probability distribution for outputs that they are sampling from. That just isn't being used for that purpose.
Ontario auditors find doctors' AI note takers routinely blow basic facts
41–50 of 141 posts
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#42I have generally moved from bearish to bullish on the future of current AI technology, but the continued inaccuracy with basic facts all while the models significantly improve continues to give me significant pause. As an example, creating recipes with Claude Opus based on flavor profiles and preferences feels magical, right up until the point at which it can't accurately convert between tablespoons and teaspoons. It…
I hate to help provide possible soultions to an entire process I don't approve of, but maybe the fuzzy tools need old style deterministic tools the same way and for the same reasons we do. So instead of an LLM trying to answer a math or reason question by finding a statistical match with other similar groups of words it found on 4chan and the all in podcast and a terrible recipe for soup written by a terrible cook, i…
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#43Earlier quoted context omitted.
could you please tell me how it generates that certainty score?
The whole thing is a statistical model, that's just what it is. No, I cannot in a reasonable way dissect how an LLM works to a satisfactory level to a skeptic.
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#44Earlier quoted context omitted.
No, they just need to be trained to have adversarial self review "thinking" processes. You ask an LLM "What's wrong with your answer?" and you get pretty good results.
Or you get the original output result was perfect and the adversarial "rethinking" switches to an incorrect result.
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#45The AI note taker we use at work records the meeting as well, and each note it takes about the meeting has a timestamp link that takes you directly there in the recording so you can check it yourself. While I'm sure a solution like this is more complicated in a HIPPAA environment, something like this is critical for things as important as healthcare.
- some human checking all the notes by listening to the entire meeting recording (takes a lot of time and man-hours)
- attendees checking notes from memory (prone to error unless they take notes)
- attendees cross checking with their own notes (defies the point of having the AI note taker)
The reality is that AI usage is not acceptable in any form in any context where accuracy is critical, but good luck getting anyone to acknowledge that.
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#46Anecdotally, we use an LLM note-taker at work for meetings. I had to intervene recently because our CIO was VERY angry at our vendor for something they promised to do and never did. He wasn't at the meeting where the "promise" was made. I was. They never promised anything, and the discussion was significantly more nuanced than what the LLM wrote in the detailed summary. In other cases, I have seen it miss the mark wh…
She called me back later that night and we chatted for bit and then she paused and sort of uncertainly was like “So… was there something you were needing to tell me?” And I was completely baffled and was like “Uhhhh I don’t think so…?”
She then explained the notification she got about my call and apparently the LLM summary of my voicemail converted a message consisting of 75% well-meaning but insignificant interpersonal human filler (like most voicemails) into this stilted, overly formal business-y speak with a somewhat ominous tone. Assigning way too much significance to each of the individual statements in the message about wanting to talk (to say happy Mother’s Day), inquiring about her availability ASAP (to say happy Mother’s Day) etc. Plus grossly exaggerating the information density of the call making it sound like I left this rambling, detailed message about needing to tell her something that was left completely vague, but possibly important and also time critical.
Added up it made her a little worried when she read it and made me a bit pissed that was the end result of my wishing her well. Because apparently everything needs a half baked LLM summary crammed into it now.
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#47Earlier quoted context omitted.
> we're not actually on the right track to achieve real intelligence. Real intelligence means you have to say "I don't know" when you don't know, or ask for help, or even just saying you refuse to help with the subtext being you don't want to appear stupid. The models could ostensibly do this when it has low confidence in it's own results but they don't. What I don't know if it's because it would be very computationa…
You can just tell the agent to do exactly that
When you point it out "Oh yes, I did do that which is contrary to the rules, request .. Anyway..."
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#48Earlier quoted context omitted.
I hate to help provide possible soultions to an entire process I don't approve of, but maybe the fuzzy tools need old style deterministic tools the same way and for the same reasons we do. So instead of an LLM trying to answer a math or reason question by finding a statistical match with other similar groups of words it found on 4chan and the all in podcast and a terrible recipe for soup written by a terrible cook, i…
I think that is how the smarter agents do things? Just like Claude/ChatGPT sometimes does a web search they can do other tool calls instead of just making a statistical guess. Of course it doesn’t always make the bright choice between those options though…
I'm available for a small fee.
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#49The AI note taker we use at work records the meeting as well, and each note it takes about the meeting has a timestamp link that takes you directly there in the recording so you can check it yourself. While I'm sure a solution like this is more complicated in a HIPPAA environment, something like this is critical for things as important as healthcare.
Re: Ontario auditors find doctors' AI note takers routinely blow basic facts
#50> 60% of evaluated AI Scribe systems mixed up prescribed drugs in patient notes, auditors say Not mentioned, as far as I can see: the comparative human mistake rate. Having seen a lot of medical records, 60% sounds about normal lol.