Live data from Hacker News

Ontario auditors find doctors' AI note takers routinely blow basic facts

theregister.com

131–140 of 141 posts

Re: Ontario auditors find doctors' AI note takers routinely blow basic facts

#131

I have generally moved from bearish to bullish on the future of current AI technology, but the continued inaccuracy with basic facts all while the models significantly improve continues to give me significant pause. As an example, creating recipes with Claude Opus based on flavor profiles and preferences feels magical, right up until the point at which it can't accurately convert between tablespoons and teaspoons. It…

If Claude occasionally overestimates the conversion, that might be an artifact of Australian tablespoons being different (4tsp/20mL vs 3tsp/15mL in the US). This error could at least be explained as a complication of the real world. (If it's saying 3.14tsp or 2tsp then I have no idea)

Usually tsp means teaspoons, and 4tsp / 20 mL vs 3tsp / 15 mL are the same ratios of tsp to mL, aren't they?

Re: Ontario auditors find doctors' AI note takers routinely blow basic facts

#132
post #86

Earlier quoted context omitted.

Hey, I'm agreement with you. I meant that these limited extracts do need to be recorded, that's all. Read the rest of the comment :)

Oops... I am deeply sorry, thank you for the heads up! It seems I've myself committed a cardinal sin that I am usually quick to point in others - rushing to reply without comprehending the full message. (Meta-oops: I realized how LLM-ish it sounds. Quick, reboot before my cover is blown!) I happen to believe that the flaw being discussed IS fundamental and inherent in the design and architecture of LLM - this is why…

No problem, this is all human and understandable. You don't sound LLM-ish, you sound like you care a lot about this.

Which both of us do.

Re: Ontario auditors find doctors' AI note takers routinely blow basic facts

#133

Earlier quoted context omitted.

> Real intelligence means you have to say "I don't know" when you don't know I have met many supposedly intelligent, certainly high status, humans who don't appear to be able to do that either. I have more confidence we can train AIs to do it, honestly.

While it is true that there are people who do not admit they are wrong when they factually are, your assertion glosses over the fact that most of the people we maintain in our social circle are people we trust through our experiences with them to be honest.

We don't always have control over who everyone we interact with though. We interact with many more people than voluntarily chosen friends.

Re: Ontario auditors find doctors' AI note takers routinely blow basic facts

#134

I have generally moved from bearish to bullish on the future of current AI technology, but the continued inaccuracy with basic facts all while the models significantly improve continues to give me significant pause. As an example, creating recipes with Claude Opus based on flavor profiles and preferences feels magical, right up until the point at which it can't accurately convert between tablespoons and teaspoons. It…

I was skeptical that LLMs could be the right path to AGI, but then I kept being impressed by how much further we could take it by expanding upon the way we use it, the harnesses we use with LLMs, and better context engineering.

When I see how LLMs are capable of essentially prompt and context engineering for themselves, it makes me think they won't need human guidance forever.

When it comes to simple fact-based tasks that have a concrete methodology, it is no surprise to me that LLMs aren't the right tool, and I believe it's a failure of the harness to not recognize those types of tasks and handle them with a more concretely functioning tool instead of relying on statistical probabilities in the LLM "brain" to spit out the correct number to a math problem.

In the same sense that LLMs can use "skills" when necessary, it should have tools or possibly even specialized "brains" for it to pass of certain types of tasks to.

I'm starting to feel that our first form of AGI is not going to be a single brain but an elaborate system of harnesses, multiple LLM models, skills, domain and task specialized subsystems it passes tasks off to, etc. Whether we get there with current LLM technology before some other evolution in AI is the question, to me.

Re: Ontario auditors find doctors' AI note takers routinely blow basic facts

#136

Can someone who is a more AI heavy user explain what is going on? I would expect an "AI Note Taker" to faithfully transcribe the entire conversation. With the same quality I see in a lot of automated video subtitles.. ie they use the wrong word a lot but it's easy to tell what they mean by context. Are these tools instead immediately summarising the whole thing, and that summary is the artifact? Because that is a bey…

Modern transformer-based STT architectures are complex but many are abstractly not entirely unlike putting the results of standard SST through an LLM with the prompt "clean this up & make it make sense". The behavior is trained in rather than prompted but the result is similar.

Obviously this results in hallucinations, mistaken implications, & inaccurately assumed context.

Re: Ontario auditors find doctors' AI note takers routinely blow basic facts

#137
But how accurate are humans? I just picked up a print out of medical history for the last 5 years and it was thick enough to be a book. There's no way a human is reading all of that and doing anything meaningful with it. Let an AI tool crunch on it and it will definitely get things wrong or jump to conclusions that aren't there, but it's quick and I can push back on those and then move to the correct answer far quicker than any meeting with a nurse or doctor will show any results. We need to focus on how to use these tools and push back on the parts that seem out of place or wrong, so we can do more rather than point out what's not perfect.

Re: Ontario auditors find doctors' AI note takers routinely blow basic facts

#138

The linked report seems almost useless -- it doesn't say anything about an error rate or a sample size, so it's a mystery whether 9 out of 20 systems “fabricated information and made suggestions to patients' treatment plans” ten out of ten times, or one out of a thousand times. If we just postulate that the systems have a high error rate, I wonder why they are being adopted. They seem extremely easy to test, so I don…

>If we just postulate that the systems have a high error rate, I wonder why they are being adopted. From the article: "While 30 percent of a platform’s evaluation score depended solely on whether they had a domestic presence in Ontario, the accuracy of medical notes contributed only 4 percent to the total score." Accuracy wasn't really part of the scoring, Ontario doesn't care about it.

Scoring systems that function by adding up several parts never make sense. Video game magazines used to do that, but it meant that you could have wretched gameplay, and still get a decent score, from points in other categories like audio, graphics, and cinematics.

Re: Ontario auditors find doctors' AI note takers routinely blow basic facts

#139

I have generally moved from bearish to bullish on the future of current AI technology, but the continued inaccuracy with basic facts all while the models significantly improve continues to give me significant pause. As an example, creating recipes with Claude Opus based on flavor profiles and preferences feels magical, right up until the point at which it can't accurately convert between tablespoons and teaspoons. It…

I was skeptical that LLMs could be the right path to AGI, but then I kept being impressed by how much further we could take it by expanding upon the way we use it, the harnesses we use with LLMs, and better context engineering. When I see how LLMs are capable of essentially prompt and context engineering for themselves, it makes me think they won't need human guidance forever. When it comes to simple fact-based tasks…

This sounds a lot like ignoring the Bitter Lesson, and expending a lot of effort rebuilding slightly better Expert Systems.

Re: Ontario auditors find doctors' AI note takers routinely blow basic facts

#140
post #37

Earlier quoted context omitted.

> It is pretty clear that initial accuracy issues will become less and less of a problem as these technologies mature. What do you base this on? As someone who can both see the amazing things genAI can do, and who sees how utterly flawed most genAI output is, it's not obvious to me. I'm working with Claude every day, Opus 4.7, and reviewing a steady stream of PRs from coworkers who are all-in, not just using due to c…

> That is a vintage hallucination that could've come right out of GPT 2.0. That's because, despite the many claims to the contrary, the models haven't actually gotten any smarter. They are still just token prediction engines at the end of the day, without any understanding of what they are doing. That's why one should not rely on them.

That is how it looks to me, too.

I'm not sure, but it seems to me that if scale or small architectural tweaks were going to solve comprehension, they would have done so by now.

Post reply on HN